Machine learning
Machine learning
The general machine-learning papers the letter carried in 2023-25. 836 featured so far, newest first.
- Featured
- 836
- Tracked on Semantic Scholar
- 833
- Cited 100+
- 195
- Since
- 16 Oct 2023
- 1 Dec 20252cites
Fairness in Survival Analysis: A Novel Conditional Mutual Information Augmentation Approach
The article introduces equalized odds in survival analysis, using Conditional Mutual Information Augmentation to improve fairness and prediction accuracy in various fields.
Machine learningOther
- 1 Dec 20250cites
A comparison of translation performance between DeepL and Supertext
The study compares DeepL and Supertext machine translation systems, finding Supertext excels at translating longer texts while emphasizing the need for context-sensitive evaluations.
Machine learningOtherIn Machine Translation Summit
- 4 Nov 20252cites
Revisiting Expected Possession Value in Football: Introducing a Benchmark, U-Net Architecture, and Reward and Risk for Passes
OJN-Pass-EPV: a new benchmark and U-Net EPV model (predicting ball height and pass risk/reward) that correctly identifies the higher-value game state about 78% of the time.
Machine learningOtherFeatured 2×
- 4 Nov 202530cites
Open Materials Generation with Stochastic Interpolants
Generative Model for Crystal Discovery: OMatG: a generative framework using stochastic interpolants and symmetry-aware (equivariant) crystal representations to design stable inorganic crystals, setting a new state of the art.
Machine learningML & AI MethodsIn International Conference on Machine Learning
- 12 Aug 202515cites
Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization
The study explores memorization in denoising score matching, revealing a regularization mechanism driven by large learning rates that prevents excessive closeness to the empirical optimal score, thus reducing memorization.
Machine learningML & AI MethodsIn Annual Conference Computational Learning Theory
- 25 Jul 20256cites
An Algebraically Converging Stochastic Gradient Descent Algorithm for Global Optimization
A new gradient descent algorithm with adaptive randomness is proposed for global optimization of nonconvex problems, proving its effectiveness and stability with numerical examples.
Machine learningOtherIn Communications in Mathematical SciencesFeatured 4×
- 25 Jul 202517cites
Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer
The paper explores the dynamics of gradient descent in deep linear networks, discussing the impact of network width and depth, and comparing various training dynamics.
Machine learningOtherIn International Conference on Machine LearningFeatured 5×
- 3 Jul 202524cites
Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation
Mosaic3D, a new data generation and training framework, has been introduced for understanding 3D scenes, achieving top results in 3D semantic and instance segmentation tasks.
Machine learningOtherIn 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)Featured 6×
- 11 Jun 202510cites
Unanswerability Evaluation for Retrieval Augmented Generation
The article introduces UAEval4RAG, a framework for evaluating the ability of retrieval-augmented generation (RAG) systems to handle unanswerable queries, emphasizing the role of component selection and prompt design.
Machine learningLLMs & TextIn Annual Meeting of the Association for Computational LinguisticsFeatured 8×
- 30 Apr 202510cites
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training
The research explores ways to speed up small language models, discovering that layer-wise adaptive pruning (Adapt-Pruner) is effective in large language models and outperforms existing pruning methods.
Machine learningLLMs & Text
- 30 Apr 20256cites
Schema-Guided Scene-Graph Reasoning Based on Multi-Agent Large Language Model System
The paper introduces SG-RwR, a new framework for reasoning and planning with scene graphs, using two large language model agents to generate task plans and information queries.
Machine learningLLMs & TextIn AAAI Conference on Artificial Intelligence
- 30 Apr 20256cites
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
The study presents SKI models, which incorporate 3D skeletons into the vision-language embedding space, using a skeleton-language model to enhance Vision Language Models and Large Vision Language Models.
Machine learningLLMs & TextIn AAAI Conference on Artificial Intelligence
- 30 Apr 20259cites
Kineto-Dynamical Planning and Accurate Execution of Minimum-Time Maneuvers on Three-Dimensional Circuits
The article introduces an artificial race driver (ARD) that learns vehicle dynamics and performs minimum-time maneuvers on a 3D track, using a new vehicle model for trajectory planning with economic nonlinear model predictive control.
Machine learningTrading, Microstructure & ExecutionIn 2025 IEEE International Conference on Robotics and Automation (ICRA)
- 23 Apr 20252cites
AAD-DCE: An Aggregated Multimodal Attention Mechanism for Early and Late Dynamic Contrast Enhanced Prostate MRI Synthesis
Multimodal Attention for MRI Synthesis: The study proposes AAD-DCE, a generative adversarial network for creating Dynamic Contrast-Enhanced MRI images, showing its superior performance compared to other DCE-MRI synthesis methods.
Machine learningML & AI MethodsIn ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)Featured 3×
- 23 Apr 20251cites
Are Language Models Up to Sequential Optimization Problems? From Evaluation to a Hegelian-Inspired Enhancement
The paper investigates the ability of Large Language Models in managing Sequential Optimization Problems, introducing WorldGen for generating new SOPs, and suggesting ACE to enhance LLM performance without additional training.
Machine learningLLMs & Text
- 9 Apr 202554cites
Do Large Language Model Benchmarks Test Reliability?
The article highlights the need for reliable large language models, criticizes current benchmarks for their inadequacy, and suggests the use of platinum benchmarks to reduce label errors and ambiguity.
Machine learningLLMs & TextFeatured 11×
- 9 Apr 202594cites
Masked Autoencoders Are Effective Tokenizers for Diffusion Models
Tokenizers for Diffusion Models: The study presents MAETok, an autoencoder for latent diffusion models, which enhances the quality of high-resolution image synthesis by learning a semantically rich latent space.
Machine learningML & AI MethodsIn International Conference on Machine LearningFeatured 11×
- 9 Apr 20258cites
Seeing World Dynamics in a Nutshell
Representing Monocular Videos Efficiently: The paper unveils NutWorld, a system that converts monocular videos into dynamic 3D Gaussian representations, offering high-quality video reconstruction and facilitating real-time applications.
Machine learningOtherFeatured 11×
- 9 Apr 20256cites
An Algebraically Converging Stochastic Gradient Descent Algorithm for Global Optimization
The authors introduce a novel gradient descent algorithm with adaptive randomness, which improves convergence rates and robustness when solving nonconvex optimization problems.
Machine learningOtherIn Communications in Mathematical SciencesFeatured 4×
- 9 Apr 202543cites
Brief analysis of DeepSeek R1 and its implications for Generative AI
Generative AI Implications: The report covers the launch of DeepSeek's new reasoning model, DeepSeekR1, its technical progress, and its impact on Generative AI, despite the US's GPU export ban.
Machine learningML & AI MethodsIn RoboticsFeatured 8×
- 9 Apr 202592cites
BFS-Prover: Scalable Best-First Tree Search for LLM-based Automatic Theorem Proving
Scalable Best-First Tree Search: BFS-Prover is a scalable framework that uses Best-First Tree Search for automatic theorem proving, challenging the need for complex tree search methods.
Machine learningLLMs & TextIn Annual Meeting of the Association for Computational LinguisticsFeatured 7×
- 9 Apr 202511cites
Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation
Python Toolkit for Retrieval and Generation: Rankify is an open-source toolkit designed to unify retrieval, re-ranking, and retrieval-augmented generation, improving consistency and scalability in information retrieval research.
Machine learningLLMs & TextFeatured 7×
- 9 Apr 202521cites
ToddlerBot: Open-Source ML-Compatible Humanoid Platform for Loco-Manipulation
Open-Source Humanoid Platform for Loco-Manipulation: ToddlerBot is a low-cost, open-source humanoid robot platform for scalable policy learning and research in robotics and AI, enabling zero-shot policy transfer.
Machine learningML & AI MethodsFeatured 7×
- 9 Apr 202557cites
NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild
Unsupervised Learning of Browser Agents: NNetNav is a method for unsupervised interaction with websites, generating synthetic demonstrations for training browser agents and making the search more tractable.
Machine learningML & AI MethodsFeatured 5×
- 9 Apr 202528cites
Dress-1-to-3: Single Image to Simulation-Ready 3D Outfit with Diffusion Prior and Differentiable Physics
Simulation-Ready 3D Outfit Generation: Dress-1-to-3 is a pipeline that reconstructs physics-plausible, simulation-ready garments and humans from an image, improving the geometric alignment of the reconstructed 3D garments and humans.
Machine learningOtherIn ACM Transactions on Graphics (TOG)Featured 2×
- 2 Apr 202519cites
QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search
Language Agents Search: QLASS system enhances the efficiency of language agents by offering step-by-step guidance, improving decision-making in complex tasks.
Machine learningML & AI MethodsIn International Conference on Machine LearningFeatured 19×
- 2 Apr 202546cites
Decision Theoretic Foundations for Conformal Prediction: Optimal Uncertainty Quantification for Risk-Averse Agents
The RAC algorithm improves decision-making in risk-sensitive areas like medicine by linking prediction uncertainty with risk-averse decision-making.
Machine learningOtherIn International Conference on Machine LearningFeatured 16×
- 2 Apr 202511cites
LoRA-X: Bridging Foundation Models with Training-Free Cross-Model Adaptation
Model Adaptation: LoRA-X enables the transfer of fine-tuning parameters across different models, enhancing the efficiency of text-to-image generation without needing original training data.
Machine learningOtherIn International Conference on Learning RepresentationsFeatured 10×
- 2 Apr 202553cites
Articulate AnyMesh: Open-Vocabulary 3D Articulated Objects Modeling
3D Object Modeling: Articulate Anymesh is a framework that transforms any rigid 3D mesh into an articulated object, aiding in the acquisition of new object manipulation skills in robotics.
Machine learningCorporate FinanceFeatured 9×
- 2 Apr 20259cites
Particle trajectory representation learning with masked point modeling
PoLAr-MAE is a self-supervised learning framework for 3D particle trajectory analysis in Time Projection Chambers, matching the performance of supervised baselines without labeled data.
Machine learningML & AI MethodsIn Machine Learning: Science and TechnologyFeatured 11×
- 2 Apr 20251cites
Hierarchical sparse Bayesian multitask learning for disease prediction in pooled microbiome studies
The article discusses a hierarchical Bayesian multitask learning model for binary classification learning, which effectively predicts human health status using microbiome profiles.
Machine learningEconometrics & ForecastingIn BioData MiningFeatured 9×
- 2 Apr 202520cites
Learning the RoPEs: Better 2D and 3D Position Encodings with STRING
The article introduces STRING, an extension of Rotary Position Encodings, which offers exact translation invariance and low computational footprint, proving beneficial in robotics and Vision Transformers.
Machine learningML & AI MethodsIn International Conference on Machine LearningFeatured 12×
- 2 Apr 202514cites
COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation
Panoptic Segmentation and Grounded Captions: The article introduces the COCONut-PanCap dataset, which improves panoptic segmentation and grounded image captioning, enhancing performance in understanding and generation tasks.
Machine learningOtherIn Neural Information Processing SystemsFeatured 9×
- 2 Apr 202544cites
Calibrated Multi-Preference Optimization for Aligning Diffusion Models
The article presents Calibrated Preference Optimization (CaPO), a method for aligning text-to-image diffusion models without human annotated data, outperforming previous methods.
Machine learningML & AI MethodsIn 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)Featured 3×
- 2 Apr 202570cites
SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration
Video Restoration with Diffusion Transformer: The article introduces SeedVR, a diffusion transformer for video restoration of any length and resolution, showing superior performance over existing methods for generic video restoration.
Machine learningML & AI MethodsIn 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)Featured 7×
- 5 Feb 202512cites
R.I.P.: Better Models by Survival of the Fittest Prompts
The study introduces Rejecting Instruction Preferences (RIP), a method for evaluating data integrity that can filter prompts or create synthetic datasets, enhancing performance across various benchmarks.
Machine learningOtherIn International Conference on Machine Learning
- 5 Feb 20251,462cites
s1: Simple test-time scaling
The research presents a method called budget forcing, which uses a small dataset to achieve test-time scaling and improved reasoning performance in language modeling, particularly in competition math questions.
Machine learningLLMs & TextIn Conference on Empirical Methods in Natural Language Processing
- 5 Feb 202547cites
Diverse Preference Optimization
The paper introduces Diverse Preference Optimization (DivPO), an optimization method that generates diverse responses in language models post-training, enhancing diversity in persona attributes and story generation.
Machine learningLLMs & Text
- 5 Feb 202547cites
Scalable-Softmax Is Superior for Attention
The study proposes Scalable-Softmax (SSMax), a replacement for Softmax in language models, which improves performance in long contexts and key information retrieval, and allows better focus on key information.
Machine learningLLMs & Text
- 5 Feb 2025171cites
Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
The research identifies underthinking in large language models, where models frequently switch reasoning thoughts, and proposes a decoding strategy to encourage deeper exploration of each reasoning path, improving accuracy across challenging datasets.
Machine learningLLMs & Text
- 5 Feb 202535cites
o3-mini vs DeepSeek-R1: Which One is Safer?
DeepSeek-R1 vs o3-mini: The AI model DeepSeek-R1 has been found to produce more unsafe responses than OpenAI's o3-mini, according to a technical report using the ASTRAL testing tool.
Machine learningML & AI Methods
- 5 Feb 20256cites
What is causal about causal models and representations?
A study presents a new framework for interpreting actions in causal Bayesian networks, addressing the limitations of current methods and enhancing the understanding of causal representation learning.
Machine learningEconometrics & Forecasting
- 5 Feb 202521cites
Prediction-Powered Inference with Imputed Covariates and Nonuniform Sampling
A novel method has been introduced to provide valid confidence intervals when machine learning algorithms fill in missing variables, extending its use to nonuniform samples and various feature subsets.
Machine learningML & AI Methods
- 5 Feb 202511cites
Decoding-based Regression
Research indicates that language models capable of numeric predictions as decoded strings perform as well as traditional methods for tabular regression tasks.
Machine learningLLMs & TextIn Trans. Mach. Learn. Res.
- 5 Feb 202589cites
SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
Visual Foundation Model for Robotic Manipulation: The new robotic manipulation system, SAM2Act, shows top-tier performance in various environments, and its memory-based version, SAM2Act+, surpasses existing methods in memory-dependent tasks.
Machine learningOtherIn International Conference on Machine Learning
- 5 Feb 202528cites
LLMs Are In-Context Bandit Reinforcement Learners
The research investigates the use of Large Language Models in in-context reinforcement learning, showing their effectiveness in learning from rewards but also their limitations in error reasoning.
Machine learningLLMs & Text
- 5 Feb 2025888cites
TÜLU 3: Pushing Frontiers in Open Language Model Post-Training
Open Language Model Post-Training: The Tulu 3 model, a top-tier post-trained language model, is introduced, outperforming other models and providing a detailed guide for its use and adaptation.
Machine learningLLMs & Text
- 5 Feb 2025223cites
SOAP: Improving and Stabilizing Shampoo using Adam
A new algorithm, SOAP, enhances the computational efficiency of the Shampoo preconditioning method in deep learning tasks, reducing iterations and time, with an online implementation available.
Machine learningML & AI Methods
- 5 Feb 202563cites
Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate
The article introduces Critique Fine-Tuning (CFT), a new method for training language models that critiques incorrect responses, showing better results than the traditional Supervised Fine-Tuning (SFT) method in math benchmarks.
Machine learningML & AI Methods
- 5 Feb 20250cites
Brain-Inspired AI with Hyperbolic Geometry
The paper suggests that using hyperbolic geometry in artificial neural networks (ANNs) and machine learning, inspired by the human brain's structure, could improve accuracy and efficiency in various tasks.
Machine learningML & AI Methods
- 5 Feb 202561cites
Optimizing Large Language Model Training Using FP4 Quantization
The research presents the first FP4 training framework for large language models (LLMs), using low-bit arithmetic operations to lessen computational demands, achieving similar accuracy to BF16 and FP8 with slight degradation.
Machine learningLLMs & TextIn International Conference on Machine Learning
- 23 Jan 20258cites
Physics of Skill Learning
The study proposes three models - Geometry, Resource, and Domino - to understand the physics of skill learning in neural networks, offering insights into neural scaling laws and learning dynamics.
Machine learningML & AI Methods
- 23 Jan 20258cites
GPS as a Control Signal for Image Generation
The research uses GPS tags in photo metadata to train models that generate images based on location, improving the estimated 3D structure and capturing the unique appearance of different locations.
Machine learningOtherIn 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
- 23 Jan 2025593cites
Continuous 3D Perception Model with Persistent State
The paper presents CUT3R, a unified framework that uses a recurrent model to generate metric-scale pointmaps from a stream of images, enabling dense scene reconstruction that updates with new images.
Machine learningOtherIn 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
- 23 Jan 202514cites
Learning Segmentation from Point Trajectories
The study introduces a method for segmenting objects in videos based on motion, using long-term point trajectories to complement optical flow, improving motion-based segmentation.
Machine learningML & AI MethodsIn Neural Information Processing Systems
- 23 Jan 202522cites
Zero-Shot Monocular Scene Flow Estimation in the Wild
The research proposes a method for scene flow prediction that estimates geometry and motion, offers a solution to scene flow data scarcity, and introduces a natural parameterization for scene flow prediction, enhancing scene flow prediction in-the-wild.
Machine learningOtherIn 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
- 23 Jan 20257cites
Expertise elevates AI usage: experimental evidence comparing laypeople and professional artists
A study shows that while AI tools can assist in artistic creation, professional artists still produce more creative and accurate work, though the difference is slight.
Machine learningML & AI MethodsIn International Journal of Human-Computer Interaction
- 23 Jan 202517cites
GauSTAR: Gaussian Surface Tracking and Reconstruction
GSTAR, a new method for photo-realistic rendering and 3D tracking of dynamic scenes, has been introduced, enabling a variety of applications.
Machine learningOtherIn 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
- 23 Jan 202555cites
DexForce: Extracting Force-Informed Actions From Kinesthetic Demonstrations for Dexterous Manipulation
DexForce, a new method for capturing demonstrations of complex manipulation, uses contact forces to compute actions for policy learning, achieving a 76% success rate.
Machine learningML & AI MethodsIn IEEE Robotics and Automation Letters
- 23 Jan 20252cites
Efficient Algorithm for Sparse Fourier Transform of Generalized q-ary Functions
GFast, a new algorithm, efficiently calculates the Fourier transform of functions over generalized q-ary sequences, outperforming existing algorithms in speed and sample usage.
Machine learningOtherIn 2025 IEEE Information Theory Workshop (ITW)
- 23 Jan 202570cites
HAC++: Towards 100X Compression of 3D Gaussian Splatting
HAC++, a new 3D Gaussian Splatting compression technique, uses relationships between unorganized anchors and a structured hash grid to achieve a size reduction of over 100X while improving fidelity.
Machine learningOtherIn IEEE Transactions on Pattern Analysis and Machine Intelligence
- 23 Jan 2025256cites
Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps
The research shows that increasing computation during inference-time can enhance the quality of samples produced by diffusion models, especially in image generation.
Machine learningML & AI Methods
- 23 Jan 2025250cites
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
The article discusses advancements in Large Language Models (LLMs) reasoning, emphasizing the use of reinforcement learning and thought simulation for complex reasoning, and the potential of scaling during training and testing.
Machine learningLLMs & Text
- 23 Jan 202526cites
OmniThink: Expanding Knowledge Boundaries in Machine Writing through Thinking
Machine Writing Expansion: OmniThink, a machine writing framework that mimics learner cognition, is introduced to improve the knowledge density of machine-written articles, addressing the limitations of retrieval-augmented generation.
Machine learningLLMs & TextIn Conference on Empirical Methods in Natural Language Processing
- 23 Jan 202533cites
Learnings from Scaling Visual Tokenizers for Reconstruction and Generation
The study reveals that scaling the decoder in auto-encoders, specifically the VisionTransformer architecture for Tokenization (ViTok), improves reconstruction performance and sets new standards for class-conditional video generation when combined with Diffusion Transformers.
Machine learningML & AI MethodsIn International Conference on Machine Learning
- 23 Jan 20253cites
Suggesting Code Edits in Interactive Machine Learning Notebooks Using Large Language Models
A study using a dataset of over 48,000 Jupyter notebook edits from GitHub reveals the complexity of machine learning maintenance tasks and the potential of large language models in predicting code edits.
Machine learningLLMs & Text
- 23 Jan 2025183cites
T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation
Text-to-Video Benchmark: TV-CompBench, a new benchmark for evaluating text-to-video generative models, shows that current models struggle with composing various elements into a video.
Machine learningML & AI MethodsIn 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
- 23 Jan 20253cites
Neuradicon: Operational representation learning of neuroimaging reports
Learning Neuroimaging Reports: Neuradicon, a new natural language processing framework, has been developed for analyzing neuroradiological reports, showing excellent adaptability across different time periods and healthcare institutions.
Machine learningML & AI MethodsIn Computer methods and programs in biomedicine
- 23 Jan 2025673cites
FAST: Efficient Action Tokenization for Vision-Language-Action Models
A new tokenization scheme, Frequency-space Action Sequence Tokenization (FAST), has been proposed for robot actions, facilitating the training of vision-language action policies for complex and high-frequency tasks.
Machine learningTrading, Microstructure & ExecutionIn Robotics
- 15 Jan 202522cites
An Empirical Study of Autoregressive Pre-Training from Videos
The study presents Toto, a series of video models trained on over 1 trillion visual tokens, showing strong performance in tasks like image recognition and object tracking.
Machine learningOtherIn 2025 IEEE/CVF International Conference on Computer Vision (ICCV)
- 15 Jan 202514cites
Decentralized Diffusion Models
The paper suggests Decentralized Diffusion Models, a framework for distributing AI model training across separate clusters, reducing costs and increasing resilience to GPU failures.
Machine learningML & AI MethodsIn 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
- 15 Jan 2025110cites
The GAN is dead; long live the GAN! A Modern GAN Baseline
The study introduces R3GAN, a simplified GAN baseline that outperforms StyleGAN2 on various datasets and competes well against other state-of-the-art GANs and diffusion models.
Machine learningML & AI Methods
- 15 Jan 202558cites
GenMol: A Drug Discovery Generalist with Discrete Diffusion
Drug Discovery Generalist: The paper presents GenMol, a molecular generative model that surpasses previous models in new generation and fragment-constrained generation, offering a unified approach for drug discovery tasks.
Machine learningML & AI MethodsIn International Conference on Machine Learning
- 15 Jan 202591cites
Neuro-Symbolic AI in 2024: A Systematic Review
Neuro-Symbolic AI has grown since 2020, focusing on learning and inference, but still lacks in areas like explainability, trustworthiness, and Meta-Cognition.
Machine learningML & AI MethodsIn LNSAI@IJCAI
- 15 Jan 202522cites
RoboPanoptes: The All-seeing Robot with Whole-body Dexterity
The All-seeing Robot: RoboPanoptes, a robot system, learns complex manipulation skills from human demonstrations using a visuomotor policy, enabling it to perform tasks like unboxing in narrow spaces and sweeping oversized objects.
Machine learningOtherIn Robotics
- 15 Jan 2025227cites
Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control
The study presents Adjoint Matching, a new algorithm that enhances dynamical generative models by refining reward fine-tuning, leading to improved consistency, realism, and adaptability to unseen human preference reward models.
Machine learningML & AI Methods
- 15 Jan 202541cites
Grokking at the Edge of Numerical Stability
The study investigates 'grokking' in deep learning, introduces Softmax Collapse and naïve loss minimization concepts, and suggests a new activation function and training algorithm for grokking without regularization.
Machine learningML & AI MethodsIn International Conference on Learning Representations
- 15 Jan 202520cites
Unity by Diversity: Improved Representation Learning in Multimodal VAEs
A new mixture-of-experts prior for Variational Autoencoders for multimodal data has been proposed, replacing hard constraints with a soft one, leading to better latent representation and improved imputation of missing data modalities.
Machine learningML & AI MethodsIn Neural Information Processing Systems
- 8 Jan 202522cites
Metadata Conditioning Accelerates Language Model Pre-training
The MeCo method speeds up language model pre-training by using additional learning cues, allowing the model to work without metadata and enhancing task performance.
Machine learningLLMs & TextIn International Conference on Machine Learning
- 8 Jan 2025230cites
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Vision and Speech Interaction: A proposed training methodology enables Large Language Models to comprehend visual and speech data, improving speech-to-speech dialogue capabilities and response speed.
Machine learningLLMs & TextIn Neural Information Processing Systems
- 8 Jan 20257cites
JOG3R: Towards 3D-Consistent Video Generators
Video Generation and Camera Pose Estimation: Research into 3D awareness in video generators shows that task-specific supervision greatly improves their accuracy for camera pose estimation.
Machine learningOther
- 8 Jan 20254cites
ProTracker: Probabilistic Integration for Robust and Accurate Point Tracking
Point Tracking: ProTracker, a new video tracking framework, combines optical flow estimations and semantic features, outperforming other unsupervised and self-supervised methods.
Machine learningOther
- 8 Jan 202510cites
VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment
3D Reconstruction from Videos: VideoLifter, a new framework, optimizes 3D representation from video sequences, speeding up the reconstruction process and surpassing other methods in visual fidelity and efficiency.
Machine learningOther
- 8 Jan 202528cites
R-SCoRe: Revisiting Scene Coordinate Regression for Robust Large-Scale Visual Localization
The study presents a new visual localization method using a covisibility graph-based global encoding learning and data augmentation strategy, achieving top results on large-scale datasets without needing network ensembles or 3D supervision.
Machine learningML & AI MethodsIn 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
- 8 Jan 202516cites
Detecting AI-Generated Text in Educational Content: Leveraging Machine Learning and Explainable AI for Academic Integrity
The research introduces tools for detecting AI-generated content in student work using machine learning and deep learning algorithms, aiming to uphold academic integrity and responsible AI use in education.
Machine learningML & AI Methods
- 8 Jan 2025395cites
Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models
The paper proposes a new model, VA-VAE, that aligns the latent space with pre-trained vision foundation models, enabling faster convergence of Diffusion Transformers in high-dimensional latent spaces and achieving top performance on ImageNet 256x256 generation.
Machine learningML & AI MethodsIn 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
- 8 Jan 202519cites
BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning
The article introduces BoostStep, a method that enhances the reasoning quality within each step of large language models solving complex math problems, providing more relevant examples and integrating seamlessly with Monte Carlo Tree Search methods.
Machine learningLLMs & Text
- 8 Jan 202524cites
Nested Attention: Semantic-aware Attention Values for Concept Personalization
The study presents Nested Attention, a mechanism that injects a rich and expressive image representation into the model's existing cross-attention layers, enabling high identity preservation while adhering to input text prompts in personalizing text-to-image models.
Machine learningOtherIn Proceedings of the Special Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers
- 8 Jan 2025140cites
MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes
The article discusses MEDEC, a benchmark for identifying and fixing medical errors in clinical notes, and reveals that while Large Language Models (LLMs) are effective, they are still not as accurate as medical doctors.
Machine learningLLMs & TextIn Annual Meeting of the Association for Computational Linguistics
- 8 Jan 2025245cites
Accurate RNA 3D structure prediction using a language model-based deep learning approach
The paper introduces RhoFold+, a deep learning method that accurately predicts 3D structures of single-chain RNAs from sequences, surpassing existing methods and aiding in RNA structure and function research.
Machine learningLLMs & TextIn Nature Methods
- 8 Jan 20255cites
FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment
Sound Effects Tool: Stable-V2A is a two-stage model that automates repetitive tasks in audio creation for video scenes, aiding sound designers in focusing on creative aspects.
Machine learningOther
- 8 Jan 202515cites
Constrained Sampling with Primal-Dual Langevin Monte Carlo
The study presents a PD-LMC algorithm that samples from a probability distribution while meeting statistical constraints, useful in Bayesian inference and prediction fairness.
Machine learningEconometrics & ForecastingIn Neural Information Processing Systems
- 8 Jan 202527cites
EdgeRAG: Online-Indexed RAG for Edge Devices
Online RAG: EdgeRAG is a system proposed for deploying Retrieval Augmented Generation on devices with limited resources, reducing latency and memory usage by pruning and generating embeddings as needed.
Machine learningLLMs & Text
- 1 Jan 202535cites
InfAlign: Inference-aware language model alignment
The study introduces a new framework for language models that enhances inference-time decoding procedures, leading to significant improvements over previous methods.
Machine learningLLMs & TextIn International Conference on Machine Learning
- 1 Jan 20250cites
Machine Learning for Sentiment Analysis of Imported Food in Trinidad and Tobago
The research shows that the VADER machine learning algorithm performs best in sentiment analysis of Twitter data on imported food in Trinidad and Tobago.
Machine learningLLMs & Text
- 1 Jan 202511cites
Symbolic approximations to Ricci-flat metrics via extrinsic symmetries of Calabi–Yau hypersurfaces
The paper uses machine learning to explore flat metrics of Fermat Calabi-Yau n-folds, revealing new properties and achieving significant reductions in Ricci curvature.
Machine learningML & AI MethodsIn Machine Learning: Science and Technology
- 1 Jan 20252cites
IMAGINE: An 8-to-1b 22nm FD-SOI Compute-In-Memory CNN Accelerator With an End-to-End Analog Charge-Based 0.15-8POPS/W Macro Featuring Distribution-Aware Data Reshaping
The paper introduces IMAGINE, a compute-in-memory SRAM for processing convolutional neural networks, which offers high energy efficiency and competitive accuracies on MNIST and CIFAR-10.
Machine learningML & AI MethodsIn IEEE Transactions on Circuits and Systems for Artificial Intelligence
- 1 Jan 20255cites
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
The article introduces a new adaptive batch size schedule for large-scale model training, which optimizes memory usage and performs better than constant batch sizes, especially in pretraining smaller models.
Machine learningLLMs & TextIn CPAL
- 1 Jan 20252cites
Tensor Network Estimation of Distribution Algorithms
The study explores the use of tensor networks in evolutionary optimization algorithms, concluding that better generative models don't always improve optimization performance and suggests adding a mutation operator for better results.
Machine learningML & AI MethodsIn Neurocomputing
- 1 Jan 20250cites
A new approach to locally adaptive polynomial regression
Locally Adaptive Nonparametric Regression: The paper presents LASER, a new nonparametric regression method that adapts to the local Hölder exponent of the regression function, outperforming other locally adaptive methods in various experiments.
Machine learningOther