---
title: Machine learning papers featured by ML-Quant
url: https://www.ml-quant.com/papers/ml/
site: ML-Quant (https://www.ml-quant.com)
updated: 2026-09-26
license: Summaries CC BY 4.0; links go to the original sources
index: https://www.ml-quant.com/llms.txt
---


# Machine learning

The general machine-learning papers the letter carried in 2023-25.

- [Fairness in Survival Analysis: A Novel Conditional Mutual Information Augmentation Approach](https://www.ml-quant.com/papers/arxiv/2502.02567/) (2025-12-01): The article introduces equalized odds in survival analysis, using Conditional Mutual Information Augmentation to improve fairness and prediction accuracy in various fields.
- [A comparison of translation performance between DeepL and Supertext](https://www.ml-quant.com/papers/arxiv/2502.02577/) (2025-12-01): The study compares DeepL and Supertext machine translation systems, finding Supertext excels at translating longer texts while emphasizing the need for context-sensitive evaluations.
- [Revisiting Expected Possession Value in Football: Introducing a Benchmark, U-Net Architecture, and Reward and Risk for Passes](https://www.ml-quant.com/papers/arxiv/2502.02565/) (2025-11-04): OJN-Pass-EPV: a new benchmark and U-Net EPV model (predicting ball height and pass risk/reward) that correctly identifies the higher-value game state about 78% of the time.
- [Open Materials Generation with Stochastic Interpolants](https://www.ml-quant.com/papers/arxiv/2502.02582/) (2025-11-04): Generative Model for Crystal Discovery: OMatG: a generative framework using stochastic interpolants and symmetry-aware (equivariant) crystal representations to design stable inorganic crystals, setting a new state of the art.
- [Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization](https://www.ml-quant.com/papers/arxiv/2502.03435/) (2025-08-12): The study explores memorization in denoising score matching, revealing a regularization mechanism driven by large learning rates that prevents excessive closeness to the empirical optimal score, thus reducing memorization.
- [An Algebraically Converging Stochastic Gradient Descent Algorithm for Global Optimization](https://www.ml-quant.com/papers/doi/10-4310-cms-250607105334/) (2025-07-25): A new gradient descent algorithm with adaptive randomness is proposed for global optimization of nonconvex problems, proving its effectiveness and stability with numerical examples.
- [Deep Linear Network Training Dynamics from Random Initialization: Data, Width, Depth, and Hyperparameter Transfer](https://www.ml-quant.com/papers/arxiv/2502.02531/) (2025-07-25): The paper explores the dynamics of gradient descent in deep linear networks, discussing the impact of network width and depth, and comparing various training dynamics.
- [Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation](https://www.ml-quant.com/papers/arxiv/2502.02548/) (2025-07-03): Mosaic3D, a new data generation and training framework, has been introduced for understanding 3D scenes, achieving top results in 3D semantic and instance segmentation tasks.
- [Unanswerability Evaluation for Retrieval Augmented Generation](https://www.ml-quant.com/papers/arxiv/2412.12300/) (2025-06-11): The article introduces UAEval4RAG, a framework for evaluating the ability of retrieval-augmented generation (RAG) systems to handle unanswerable queries, emphasizing the role of component selection and prompt design.
- [Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training](https://www.ml-quant.com/papers/arxiv/2502.03460/) (2025-04-30): The research explores ways to speed up small language models, discovering that layer-wise adaptive pruning (Adapt-Pruner) is effective in large language models and outperforms existing pruning methods.
- [Schema-Guided Scene-Graph Reasoning Based on Multi-Agent Large Language Model System](https://www.ml-quant.com/papers/arxiv/2502.03450/) (2025-04-30): The paper introduces SG-RwR, a new framework for reasoning and planning with scene graphs, using two large language model agents to generate task plans and information queries.
- [SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living](https://www.ml-quant.com/papers/arxiv/2502.03459/) (2025-04-30): The study presents SKI models, which incorporate 3D skeletons into the vision-language embedding space, using a skeleton-language model to enhance Vision Language Models and Large Vision Language Models.
- [Kineto-Dynamical Planning and Accurate Execution of Minimum-Time Maneuvers on Three-Dimensional Circuits](https://www.ml-quant.com/papers/arxiv/2502.03454/) (2025-04-30): The article introduces an artificial race driver (ARD) that learns vehicle dynamics and performs minimum-time maneuvers on a 3D track, using a new vehicle model for trajectory planning with economic nonlinear model predictive control.
- [AAD-DCE: An Aggregated Multimodal Attention Mechanism for Early and Late Dynamic Contrast Enhanced Prostate MRI Synthesis](https://www.ml-quant.com/papers/arxiv/2502.02555/) (2025-04-23): Multimodal Attention for MRI Synthesis: The study proposes AAD-DCE, a generative adversarial network for creating Dynamic Contrast-Enhanced MRI images, showing its superior performance compared to other DCE-MRI synthesis methods.
- [Are Language Models Up to Sequential Optimization Problems? From Evaluation to a Hegelian-Inspired Enhancement](https://www.ml-quant.com/papers/arxiv/2502.02573/) (2025-04-23): The paper investigates the ability of Large Language Models in managing Sequential Optimization Problems, introducing WorldGen for generating new SOPs, and suggesting ACE to enhance LLM performance without additional training.
- [Do Large Language Model Benchmarks Test Reliability?](https://www.ml-quant.com/papers/arxiv/2502.03461/) (2025-04-09): The article highlights the need for reliable large language models, criticizes current benchmarks for their inadequacy, and suggests the use of platinum benchmarks to reduce label errors and ambiguity.
- [Masked Autoencoders Are Effective Tokenizers for Diffusion Models](https://www.ml-quant.com/papers/arxiv/2502.03444/) (2025-04-09): Tokenizers for Diffusion Models: The study presents MAETok, an autoencoder for latent diffusion models, which enhances the quality of high-resolution image synthesis by learning a semantically rich latent space.
- [Seeing World Dynamics in a Nutshell](https://www.ml-quant.com/papers/arxiv/2502.03465/) (2025-04-09): Representing Monocular Videos Efficiently: The paper unveils NutWorld, a system that converts monocular videos into dynamic 3D Gaussian representations, offering high-quality video reconstruction and facilitating real-time applications.
- [An Algebraically Converging Stochastic Gradient Descent Algorithm for Global Optimization](https://www.ml-quant.com/papers/arxiv/2204.05923/) (2025-04-09): The authors introduce a novel gradient descent algorithm with adaptive randomness, which improves convergence rates and robustness when solving nonconvex optimization problems.
- [Brief analysis of DeepSeek R1 and its implications for Generative AI](https://www.ml-quant.com/papers/arxiv/2502.02523/) (2025-04-09): Generative AI Implications: The report covers the launch of DeepSeek's new reasoning model, DeepSeekR1, its technical progress, and its impact on Generative AI, despite the US's GPU export ban.
- [BFS-Prover: Scalable Best-First Tree Search for LLM-based Automatic Theorem Proving](https://www.ml-quant.com/papers/arxiv/2502.03438/) (2025-04-09): Scalable Best-First Tree Search: BFS-Prover is a scalable framework that uses Best-First Tree Search for automatic theorem proving, challenging the need for complex tree search methods.
- [Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation](https://www.ml-quant.com/papers/arxiv/2502.02464/) (2025-04-09): Python Toolkit for Retrieval and Generation: Rankify is an open-source toolkit designed to unify retrieval, re-ranking, and retrieval-augmented generation, improving consistency and scalability in information retrieval research.
- [ToddlerBot: Open-Source ML-Compatible Humanoid Platform for Loco-Manipulation](https://www.ml-quant.com/papers/arxiv/2502.00893/) (2025-04-09): Open-Source Humanoid Platform for Loco-Manipulation: ToddlerBot is a low-cost, open-source humanoid robot platform for scalable policy learning and research in robotics and AI, enabling zero-shot policy transfer.
- [NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild](https://www.ml-quant.com/papers/arxiv/2410.02907/) (2025-04-09): Unsupervised Learning of Browser Agents: NNetNav is a method for unsupervised interaction with websites, generating synthetic demonstrations for training browser agents and making the search more tractable.
- [Dress-1-to-3: Single Image to Simulation-Ready 3D Outfit with Diffusion Prior and Differentiable Physics](https://www.ml-quant.com/papers/arxiv/2502.03449/) (2025-04-09): Simulation-Ready 3D Outfit Generation: Dress-1-to-3 is a pipeline that reconstructs physics-plausible, simulation-ready garments and humans from an image, improving the geometric alignment of the reconstructed 3D garments and humans.
- [QLASS: Boosting Language Agent Inference via Q-Guided Stepwise Search](https://www.ml-quant.com/papers/arxiv/2502.02584/) (2025-04-02): Language Agents Search: QLASS system enhances the efficiency of language agents by offering step-by-step guidance, improving decision-making in complex tasks.
- [Decision Theoretic Foundations for Conformal Prediction: Optimal Uncertainty Quantification for Risk-Averse Agents](https://www.ml-quant.com/papers/arxiv/2502.02561/) (2025-04-02): The RAC algorithm improves decision-making in risk-sensitive areas like medicine by linking prediction uncertainty with risk-averse decision-making.
- [LoRA-X: Bridging Foundation Models with Training-Free Cross-Model Adaptation](https://www.ml-quant.com/papers/arxiv/2501.16559/) (2025-04-02): Model Adaptation: LoRA-X enables the transfer of fine-tuning parameters across different models, enhancing the efficiency of text-to-image generation without needing original training data.
- [Articulate AnyMesh: Open-Vocabulary 3D Articulated Objects Modeling](https://www.ml-quant.com/papers/arxiv/2502.02590/) (2025-04-02): 3D Object Modeling: Articulate Anymesh is a framework that transforms any rigid 3D mesh into an articulated object, aiding in the acquisition of new object manipulation skills in robotics.
- [Particle trajectory representation learning with masked point modeling](https://www.ml-quant.com/papers/arxiv/2502.02558/) (2025-04-02): PoLAr-MAE is a self-supervised learning framework for 3D particle trajectory analysis in Time Projection Chambers, matching the performance of supervised baselines without labeled data.
- [Hierarchical sparse Bayesian multitask learning for disease prediction in pooled microbiome studies](https://www.ml-quant.com/papers/arxiv/2502.02552/) (2025-04-02): The article discusses a hierarchical Bayesian multitask learning model for binary classification learning, which effectively predicts human health status using microbiome profiles.
- [Learning the RoPEs: Better 2D and 3D Position Encodings with STRING](https://www.ml-quant.com/papers/arxiv/2502.02562/) (2025-04-02): The article introduces STRING, an extension of Rotary Position Encodings, which offers exact translation invariance and low computational footprint, proving beneficial in robotics and Vision Transformers.
- [COCONut-PanCap: Joint Panoptic Segmentation and Grounded Captions for Fine-Grained Understanding and Generation](https://www.ml-quant.com/papers/arxiv/2502.02589/) (2025-04-02): Panoptic Segmentation and Grounded Captions: The article introduces the COCONut-PanCap dataset, which improves panoptic segmentation and grounded image captioning, enhancing performance in understanding and generation tasks.
- [Calibrated Multi-Preference Optimization for Aligning Diffusion Models](https://www.ml-quant.com/papers/arxiv/2502.02588/) (2025-04-02): The article presents Calibrated Preference Optimization (CaPO), a method for aligning text-to-image diffusion models without human annotated data, outperforming previous methods.
- [SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration](https://www.ml-quant.com/papers/arxiv/2501.01320/) (2025-04-02): Video Restoration with Diffusion Transformer: The article introduces SeedVR, a diffusion transformer for video restoration of any length and resolution, showing superior performance over existing methods for generic video restoration.
- [R.I.P.: Better Models by Survival of the Fittest Prompts](https://www.ml-quant.com/papers/arxiv/2501.18578/) (2025-02-05): The study introduces Rejecting Instruction Preferences (RIP), a method for evaluating data integrity that can filter prompts or create synthetic datasets, enhancing performance across various benchmarks.
- [s1: Simple test-time scaling](https://www.ml-quant.com/papers/arxiv/2501.19393/) (2025-02-05): The research presents a method called budget forcing, which uses a small dataset to achieve test-time scaling and improved reasoning performance in language modeling, particularly in competition math questions.
- [Diverse Preference Optimization](https://www.ml-quant.com/papers/arxiv/2501.18101/) (2025-02-05): The paper introduces Diverse Preference Optimization (DivPO), an optimization method that generates diverse responses in language models post-training, enhancing diversity in persona attributes and story generation.
- [Scalable-Softmax Is Superior for Attention](https://www.ml-quant.com/papers/arxiv/2501.19399/) (2025-02-05): The study proposes Scalable-Softmax (SSMax), a replacement for Softmax in language models, which improves performance in long contexts and key information retrieval, and allows better focus on key information.
- [Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs](https://www.ml-quant.com/papers/arxiv/2501.18585/) (2025-02-05): The research identifies underthinking in large language models, where models frequently switch reasoning thoughts, and proposes a decoding strategy to encourage deeper exploration of each reasoning path, improving accuracy across challenging datasets.
- [o3-mini vs DeepSeek-R1: Which One is Safer?](https://www.ml-quant.com/papers/arxiv/2501.18438/) (2025-02-05): DeepSeek-R1 vs o3-mini: The AI model DeepSeek-R1 has been found to produce more unsafe responses than OpenAI's o3-mini, according to a technical report using the ASTRAL testing tool.
- [What is causal about causal models and representations?](https://www.ml-quant.com/papers/arxiv/2501.19335/) (2025-02-05): A study presents a new framework for interpreting actions in causal Bayesian networks, addressing the limitations of current methods and enhancing the understanding of causal representation learning.
- [Prediction-Powered Inference with Imputed Covariates and Nonuniform Sampling](https://www.ml-quant.com/papers/arxiv/2501.18577/) (2025-02-05): A novel method has been introduced to provide valid confidence intervals when machine learning algorithms fill in missing variables, extending its use to nonuniform samples and various feature subsets.
- [Decoding-based Regression](https://www.ml-quant.com/papers/arxiv/2501.19383/) (2025-02-05): Research indicates that language models capable of numeric predictions as decoded strings perform as well as traditional methods for tabular regression tasks.
- [SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation](https://www.ml-quant.com/papers/arxiv/2501.18564/) (2025-02-05): Visual Foundation Model for Robotic Manipulation: The new robotic manipulation system, SAM2Act, shows top-tier performance in various environments, and its memory-based version, SAM2Act+, surpasses existing methods in memory-dependent tasks.
- [LLMs Are In-Context Bandit Reinforcement Learners](https://www.ml-quant.com/papers/arxiv/2410.05362/) (2025-02-05): The research investigates the use of Large Language Models in in-context reinforcement learning, showing their effectiveness in learning from rewards but also their limitations in error reasoning.
- [TÜLU 3: Pushing Frontiers in Open Language Model Post-Training](https://www.ml-quant.com/papers/arxiv/2411.15124/) (2025-02-05): Open Language Model Post-Training: The Tulu 3 model, a top-tier post-trained language model, is introduced, outperforming other models and providing a detailed guide for its use and adaptation.
- [SOAP: Improving and Stabilizing Shampoo using Adam](https://www.ml-quant.com/papers/arxiv/2409.11321/) (2025-02-05): A new algorithm, SOAP, enhances the computational efficiency of the Shampoo preconditioning method in deep learning tasks, reducing iterations and time, with an online implementation available.
- [Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate](https://www.ml-quant.com/papers/arxiv/2501.17703/) (2025-02-05): The article introduces Critique Fine-Tuning (CFT), a new method for training language models that critiques incorrect responses, showing better results than the traditional Supervised Fine-Tuning (SFT) method in math benchmarks.
- [Brain-Inspired AI with Hyperbolic Geometry](https://www.ml-quant.com/papers/arxiv/2409.12990/) (2025-02-05): The paper suggests that using hyperbolic geometry in artificial neural networks (ANNs) and machine learning, inspired by the human brain's structure, could improve accuracy and efficiency in various tasks.
- [Optimizing Large Language Model Training Using FP4 Quantization](https://www.ml-quant.com/papers/arxiv/2501.17116/) (2025-02-05): The research presents the first FP4 training framework for large language models (LLMs), using low-bit arithmetic operations to lessen computational demands, achieving similar accuracy to BF16 and FP8 with slight degradation.
- [Physics of Skill Learning](https://www.ml-quant.com/papers/arxiv/2501.12391/) (2025-01-23): The study proposes three models - Geometry, Resource, and Domino - to understand the physics of skill learning in neural networks, offering insights into neural scaling laws and learning dynamics.
- [GPS as a Control Signal for Image Generation](https://www.ml-quant.com/papers/arxiv/2501.12390/) (2025-01-23): The research uses GPS tags in photo metadata to train models that generate images based on location, improving the estimated 3D structure and capturing the unique appearance of different locations.
- [Continuous 3D Perception Model with Persistent State](https://www.ml-quant.com/papers/arxiv/2501.12387/) (2025-01-23): The paper presents CUT3R, a unified framework that uses a recurrent model to generate metric-scale pointmaps from a stream of images, enabling dense scene reconstruction that updates with new images.
- [Learning Segmentation from Point Trajectories](https://www.ml-quant.com/papers/arxiv/2501.12392/) (2025-01-23): The study introduces a method for segmenting objects in videos based on motion, using long-term point trajectories to complement optical flow, improving motion-based segmentation.
- [Zero-Shot Monocular Scene Flow Estimation in the Wild](https://www.ml-quant.com/papers/arxiv/2501.10357/) (2025-01-23): The research proposes a method for scene flow prediction that estimates geometry and motion, offers a solution to scene flow data scarcity, and introduces a natural parameterization for scene flow prediction, enhancing scene flow prediction in-the-wild.
- [Expertise elevates AI usage: experimental evidence comparing laypeople and professional artists](https://www.ml-quant.com/papers/arxiv/2501.12374/) (2025-01-23): A study shows that while AI tools can assist in artistic creation, professional artists still produce more creative and accurate work, though the difference is slight.
- [GauSTAR: Gaussian Surface Tracking and Reconstruction](https://www.ml-quant.com/papers/arxiv/2501.10283/) (2025-01-23): GSTAR, a new method for photo-realistic rendering and 3D tracking of dynamic scenes, has been introduced, enabling a variety of applications.
- [DexForce: Extracting Force-Informed Actions From Kinesthetic Demonstrations for Dexterous Manipulation](https://www.ml-quant.com/papers/arxiv/2501.10356/) (2025-01-23): DexForce, a new method for capturing demonstrations of complex manipulation, uses contact forces to compute actions for policy learning, achieving a 76% success rate.
- [Efficient Algorithm for Sparse Fourier Transform of Generalized q-ary Functions](https://www.ml-quant.com/papers/arxiv/2501.12365/) (2025-01-23): GFast, a new algorithm, efficiently calculates the Fourier transform of functions over generalized q-ary sequences, outperforming existing algorithms in speed and sample usage.
- [HAC++: Towards 100X Compression of 3D Gaussian Splatting](https://www.ml-quant.com/papers/arxiv/2501.12255/) (2025-01-23): HAC++, a new 3D Gaussian Splatting compression technique, uses relationships between unorganized anchors and a structured hash grid to achieve a size reduction of over 100X while improving fidelity.
- [Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps](https://www.ml-quant.com/papers/arxiv/2501.09732/) (2025-01-23): The research shows that increasing computation during inference-time can enhance the quality of samples produced by diffusion models, especially in image generation.
- [Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models](https://www.ml-quant.com/papers/arxiv/2501.09686/) (2025-01-23): The article discusses advancements in Large Language Models (LLMs) reasoning, emphasizing the use of reinforcement learning and thought simulation for complex reasoning, and the potential of scaling during training and testing.
- [OmniThink: Expanding Knowledge Boundaries in Machine Writing through Thinking](https://www.ml-quant.com/papers/arxiv/2501.09751/) (2025-01-23): Machine Writing Expansion: OmniThink, a machine writing framework that mimics learner cognition, is introduced to improve the knowledge density of machine-written articles, addressing the limitations of retrieval-augmented generation.
- [Learnings from Scaling Visual Tokenizers for Reconstruction and Generation](https://www.ml-quant.com/papers/arxiv/2501.09755/) (2025-01-23): The study reveals that scaling the decoder in auto-encoders, specifically the VisionTransformer architecture for Tokenization (ViTok), improves reconstruction performance and sets new standards for class-conditional video generation when combined with Diffusion Transformers.
- [Suggesting Code Edits in Interactive Machine Learning Notebooks Using Large Language Models](https://www.ml-quant.com/papers/arxiv/2501.09745/) (2025-01-23): A study using a dataset of over 48,000 Jupyter notebook edits from GitHub reveals the complexity of machine learning maintenance tasks and the potential of large language models in predicting code edits.
- [T2V-CompBench: A Comprehensive Benchmark for Compositional Text-to-video Generation](https://www.ml-quant.com/papers/arxiv/2407.14505/) (2025-01-23): Text-to-Video Benchmark: TV-CompBench, a new benchmark for evaluating text-to-video generative models, shows that current models struggle with composing various elements into a video.
- [Neuradicon: Operational representation learning of neuroimaging reports](https://www.ml-quant.com/papers/arxiv/2107.10021/) (2025-01-23): Learning Neuroimaging Reports: Neuradicon, a new natural language processing framework, has been developed for analyzing neuroradiological reports, showing excellent adaptability across different time periods and healthcare institutions.
- [FAST: Efficient Action Tokenization for Vision-Language-Action Models](https://www.ml-quant.com/papers/arxiv/2501.09747/) (2025-01-23): A new tokenization scheme, Frequency-space Action Sequence Tokenization (FAST), has been proposed for robot actions, facilitating the training of vision-language action policies for complex and high-frequency tasks.
- [An Empirical Study of Autoregressive Pre-Training from Videos](https://www.ml-quant.com/papers/arxiv/2501.05453/) (2025-01-15): The study presents Toto, a series of video models trained on over 1 trillion visual tokens, showing strong performance in tasks like image recognition and object tracking.
- [Decentralized Diffusion Models](https://www.ml-quant.com/papers/arxiv/2501.05450/) (2025-01-15): The paper suggests Decentralized Diffusion Models, a framework for distributing AI model training across separate clusters, reducing costs and increasing resilience to GPU failures.
- [The GAN is dead; long live the GAN! A Modern GAN Baseline](https://www.ml-quant.com/papers/arxiv/2501.05441/) (2025-01-15): The study introduces R3GAN, a simplified GAN baseline that outperforms StyleGAN2 on various datasets and competes well against other state-of-the-art GANs and diffusion models.
- [GenMol: A Drug Discovery Generalist with Discrete Diffusion](https://www.ml-quant.com/papers/arxiv/2501.06158/) (2025-01-15): Drug Discovery Generalist: The paper presents GenMol, a molecular generative model that surpasses previous models in new generation and fragment-constrained generation, offering a unified approach for drug discovery tasks.
- [Neuro-Symbolic AI in 2024: A Systematic Review](https://www.ml-quant.com/papers/arxiv/2501.05435/) (2025-01-15): Neuro-Symbolic AI has grown since 2020, focusing on learning and inference, but still lacks in areas like explainability, trustworthiness, and Meta-Cognition.
- [RoboPanoptes: The All-seeing Robot with Whole-body Dexterity](https://www.ml-quant.com/papers/arxiv/2501.05420/) (2025-01-15): The All-seeing Robot: RoboPanoptes, a robot system, learns complex manipulation skills from human demonstrations using a visuomotor policy, enabling it to perform tasks like unboxing in narrow spaces and sweeping oversized objects.
- [Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control](https://www.ml-quant.com/papers/arxiv/2409.08861/) (2025-01-15): The study presents Adjoint Matching, a new algorithm that enhances dynamical generative models by refining reward fine-tuning, leading to improved consistency, realism, and adaptability to unseen human preference reward models.
- [Grokking at the Edge of Numerical Stability](https://www.ml-quant.com/papers/arxiv/2501.04697/) (2025-01-15): The study investigates 'grokking' in deep learning, introduces Softmax Collapse and naïve loss minimization concepts, and suggests a new activation function and training algorithm for grokking without regularization.
- [Unity by Diversity: Improved Representation Learning in Multimodal VAEs](https://www.ml-quant.com/papers/arxiv/2403.05300/) (2025-01-15): A new mixture-of-experts prior for Variational Autoencoders for multimodal data has been proposed, replacing hard constraints with a soft one, leading to better latent representation and improved imputation of missing data modalities.
- [Metadata Conditioning Accelerates Language Model Pre-training](https://www.ml-quant.com/papers/arxiv/2501.01956/) (2025-01-08): The MeCo method speeds up language model pre-training by using additional learning cues, allowing the model to work without metadata and enhancing task performance.
- [VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction](https://www.ml-quant.com/papers/arxiv/2501.01957/) (2025-01-08): Vision and Speech Interaction: A proposed training methodology enables Large Language Models to comprehend visual and speech data, improving speech-to-speech dialogue capabilities and response speed.
- [JOG3R: Towards 3D-Consistent Video Generators](https://www.ml-quant.com/papers/arxiv/2501.01409/) (2025-01-08): Video Generation and Camera Pose Estimation: Research into 3D awareness in video generators shows that task-specific supervision greatly improves their accuracy for camera pose estimation.
- [ProTracker: Probabilistic Integration for Robust and Accurate Point Tracking](https://www.ml-quant.com/papers/arxiv/2501.03220/) (2025-01-08): Point Tracking: ProTracker, a new video tracking framework, combines optical flow estimations and semantic features, outperforming other unsupervised and self-supervised methods.
- [VideoLifter: Lifting Videos to 3D with Fast Hierarchical Stereo Alignment](https://www.ml-quant.com/papers/arxiv/2501.01949/) (2025-01-08): 3D Reconstruction from Videos: VideoLifter, a new framework, optimizes 3D representation from video sequences, speeding up the reconstruction process and surpassing other methods in visual fidelity and efficiency.
- [R-SCoRe: Revisiting Scene Coordinate Regression for Robust Large-Scale Visual Localization](https://www.ml-quant.com/papers/arxiv/2501.01421/) (2025-01-08): The study presents a new visual localization method using a covisibility graph-based global encoding learning and data augmentation strategy, achieving top results on large-scale datasets without needing network ensembles or 3D supervision.
- [Detecting AI-Generated Text in Educational Content: Leveraging Machine Learning and Explainable AI for Academic Integrity](https://www.ml-quant.com/papers/arxiv/2501.03203/) (2025-01-08): The research introduces tools for detecting AI-generated content in student work using machine learning and deep learning algorithms, aiming to uphold academic integrity and responsible AI use in education.
- [Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion Models](https://www.ml-quant.com/papers/arxiv/2501.01423/) (2025-01-08): The paper proposes a new model, VA-VAE, that aligns the latent space with pre-trained vision foundation models, enabling faster convergence of Diffusion Transformers in high-dimensional latent spaces and achieving top performance on ImageNet 256x256 generation.
- [BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning](https://www.ml-quant.com/papers/arxiv/2501.03226/) (2025-01-08): The article introduces BoostStep, a method that enhances the reasoning quality within each step of large language models solving complex math problems, providing more relevant examples and integrating seamlessly with Monte Carlo Tree Search methods.
- [Nested Attention: Semantic-aware Attention Values for Concept Personalization](https://www.ml-quant.com/papers/arxiv/2501.01407/) (2025-01-08): The study presents Nested Attention, a mechanism that injects a rich and expressive image representation into the model's existing cross-attention layers, enabling high identity preservation while adhering to input text prompts in personalizing text-to-image models.
- [MEDEC: A Benchmark for Medical Error Detection and Correction in Clinical Notes](https://www.ml-quant.com/papers/arxiv/2412.19260/) (2025-01-08): The article discusses MEDEC, a benchmark for identifying and fixing medical errors in clinical notes, and reveals that while Large Language Models (LLMs) are effective, they are still not as accurate as medical doctors.
- [Accurate RNA 3D structure prediction using a language model-based deep learning approach](https://www.ml-quant.com/papers/arxiv/2207.01586/) (2025-01-08): The paper introduces RhoFold+, a deep learning method that accurately predicts 3D structures of single-chain RNAs from sequences, surpassing existing methods and aiding in RNA structure and function research.
- [FolAI: Synchronized Foley Sound Generation with Semantic and Temporal Alignment](https://www.ml-quant.com/papers/arxiv/2412.15023/) (2025-01-08): Sound Effects Tool: Stable-V2A is a two-stage model that automates repetitive tasks in audio creation for video scenes, aiding sound designers in focusing on creative aspects.
- [Constrained Sampling with Primal-Dual Langevin Monte Carlo](https://www.ml-quant.com/papers/arxiv/2411.00568/) (2025-01-08): The study presents a PD-LMC algorithm that samples from a probability distribution while meeting statistical constraints, useful in Bayesian inference and prediction fairness.
- [EdgeRAG: Online-Indexed RAG for Edge Devices](https://www.ml-quant.com/papers/arxiv/2412.21023/) (2025-01-08): Online RAG: EdgeRAG is a system proposed for deploying Retrieval Augmented Generation on devices with limited resources, reducing latency and memory usage by pruning and generating embeddings as needed.
- [InfAlign: Inference-aware language model alignment](https://www.ml-quant.com/papers/arxiv/2412.19792/) (2025-01-01): The study introduces a new framework for language models that enhances inference-time decoding procedures, leading to significant improvements over previous methods.
- [Machine Learning for Sentiment Analysis of Imported Food in Trinidad and Tobago](https://www.ml-quant.com/papers/arxiv/2412.19781/) (2025-01-01): The research shows that the VADER machine learning algorithm performs best in sentiment analysis of Twitter data on imported food in Trinidad and Tobago.
- [Symbolic approximations to Ricci-flat metrics via extrinsic symmetries of Calabi–Yau hypersurfaces](https://www.ml-quant.com/papers/arxiv/2412.19778/) (2025-01-01): The paper uses machine learning to explore flat metrics of Fermat Calabi-Yau n-folds, revealing new properties and achieving significant reductions in Ricci curvature.
- [IMAGINE: An 8-to-1b 22nm FD-SOI Compute-In-Memory CNN Accelerator With an End-to-End Analog Charge-Based 0.15-8POPS/W Macro Featuring Distribution-Aware Data Reshaping](https://www.ml-quant.com/papers/arxiv/2412.19750/) (2025-01-01): The paper introduces IMAGINE, a compute-in-memory SRAM for processing convolutional neural networks, which offers high energy efficiency and competitive accuracies on MNIST and CIFAR-10.
- [Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism](https://www.ml-quant.com/papers/arxiv/2412.21124/) (2025-01-01): The article introduces a new adaptive batch size schedule for large-scale model training, which optimizes memory usage and performs better than constant batch sizes, especially in pretraining smaller models.
- [Tensor Network Estimation of Distribution Algorithms](https://www.ml-quant.com/papers/arxiv/2412.19780/) (2025-01-01): The study explores the use of tensor networks in evolutionary optimization algorithms, concluding that better generative models don't always improve optimization performance and suggests adding a mutation operator for better results.
- [A new approach to locally adaptive polynomial regression](https://www.ml-quant.com/papers/arxiv/2412.19802/) (2025-01-01): Locally Adaptive Nonparametric Regression: The paper presents LASER, a new nonparametric regression method that adapts to the local Hölder exponent of the regression function, outperforming other locally adaptive methods in various experiments.
- [Training Software Engineering Agents and Verifiers with SWE-Gym](https://www.ml-quant.com/papers/arxiv/2412.21139/) (2025-01-01): The authors introduce SWE-Gym, the first training environment for software engineering agents, featuring real-world Python tasks and showing significant improvements in task resolution rates.
- [Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?](https://www.ml-quant.com/papers/arxiv/2407.21792/) (2025-01-01): Research suggests AI safety benchmarks often align with general capabilities and training compute, leading to potential safetywashing, and recommends a stricter framework for AI safety research.
- [Long-Form Speech Generation with Spoken Language Models](https://www.ml-quant.com/papers/arxiv/2412.18603/) (2025-01-01): Google introduces SpeechSSM, a speech language model for generating long-form audio, along with new metrics and a benchmark for long-form speech processing and generation.
- [Mitigating optimistic bias in entropic risk estimation and optimization](https://www.ml-quant.com/papers/arxiv/2409.19926/) (2025-01-01): A novel bootstrapping method is suggested to reduce bias in the empirical entropic risk estimator, a tool used in high-stakes decision making, and is applied to insurance contract design.
- [Decentralized Intelligence in GameFi: Embodied AI Agents and the Convergence of DeFi and Virtual Ecosystems](https://www.ml-quant.com/papers/arxiv/2412.18601/) (2025-01-01): A proposed GameFi ecosystem integrates advanced AI agents into gaming platforms, improving player engagement and economic interaction within gaming ecosystems.
- [Local and Mixing-Based Algorithms for Gaussian Graphical Model Selection from Glauber Dynamics](https://www.ml-quant.com/papers/arxiv/2412.18594/) (2025-01-01): A new algorithm for Gaussian graphical model selection is presented, offering theoretical guarantees on computational and statistical complexity when data is sampled according to the Glauber dynamics.
- [Keypoint Aware Masked Image Modelling](https://www.ml-quant.com/papers/arxiv/2407.13873/) (2025-01-01): KAMIM is a new method that enhances vision transformers' performance by using patch-wise weighting from keypoint features, showing improved accuracy on the ImageNet-1K dataset.
- [YuLan-Mini: An Open Data-efficient Language Model](https://www.ml-quant.com/papers/arxiv/2412.17743/) (2025-01-01): Efficient Language Model: YuLan-Mini, a 2.42B parameter base model, delivers top-tier performance among similar models through a sophisticated data pipeline, robust optimization method, and effective annealing approach.
- [Can LLMs Obfuscate Code? A Systematic Analysis of Large Language Models into Assembly Code Obfuscation](https://www.ml-quant.com/papers/arxiv/2412.16135/) (2025-01-01): The MetamorphASM benchmark is designed to assess Large Language Models' ability to generate and analyze obfuscated assembly code, potentially threatening anti-virus engines.
- [Principal Component Flow Map Learning of PDEs from Incomplete, Limited, and Noisy Data](https://www.ml-quant.com/papers/arxiv/2407.10854/) (2025-01-01): A new computational technique has been introduced for modeling the evolution of dynamical systems, specifically targeting the complex problem of modeling partially-observed partial differential equations on high-dimensional non-uniform grids.
- [DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation](https://www.ml-quant.com/papers/arxiv/2412.18597/) (2025-01-01): Attention Control in Video Generation: DiTCtrl is a new training-free method for multi-prompt video generation under MM-DiT architectures, allowing for mask-guided precise semantic control across different prompts.
- [CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models](https://www.ml-quant.com/papers/arxiv/2412.12093/) (2024-12-18): Portrait Avatars: CAP4D is a method that uses a unique model to create and animate realistic 4D portrait avatars from any number of reference images in real time.
- [Representing Long Volumetric Video with Temporal Gaussian Hierarchy](https://www.ml-quant.com/papers/doi/10-1145-3687919/) (2024-12-18): The paper introduces Temporal Gaussian Hierarchy, a new 4D representation that efficiently models long volumetric videos, reducing the number of Gaussian primitives and maintaining constant GPU memory usage.
- [Stereo4D: Learning How Things Move in 3D from Internet Stereo Videos](https://www.ml-quant.com/papers/arxiv/2412.09621/) (2024-12-18): Learning Motion from Videos: The authors have developed a system that mines high-quality 4D reconstructions from internet videos, allowing for the prediction of structure and 3D motion from real-world image pairs.
- [GenEx: Generating an Explorable World](https://www.ml-quant.com/papers/arxiv/2412.09624/) (2024-12-18): Explorable World: GenEx is a system that plans complex world exploration, guided by its generative imagination about the surrounding environments, enabling AI agents to perform complex tasks.
- [Feat2GS: Probing Visual Foundation Models with Gaussian Splatting](https://www.ml-quant.com/papers/arxiv/2412.09606/) (2024-12-18): Visual Foundation Models: Feat2GS is a framework that extracts 3D Gaussians attributes from unposed images, allowing for the probing of 3D awareness for geometry and texture through novel view synthesis.
- [Causal Diffusion Transformers for Generative Modeling](https://www.ml-quant.com/papers/arxiv/2412.12095/) (2024-12-18): The article discusses Causal Diffusion, a framework that enhances diffusion models' performance and allows a seamless shift between autoregressive and diffusion generation modes, achieving top results on the ImageNet generation benchmark.
- [Spectral Image Tokenizer](https://www.ml-quant.com/papers/arxiv/2412.09607/) (2024-12-18): The paper suggests a novel method for tokenizing images for autoregressive transformer-based image generation, using a discrete wavelet transform for a coarse-to-fine representation, offering benefits like improved next-token prediction and the ability to reconstruct varying resolution images.
- [MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization](https://www.ml-quant.com/papers/arxiv/2412.12098/) (2024-12-18): The research introduces MaxInfoRL, a reinforcement learning framework that balances exploration by directing it towards informative transitions, demonstrating superior performance in challenging exploration problems and complex visual control tasks.
- [PanSplat: 4K Panorama Synthesis with Feed-Forward Gaussian Splatting](https://www.ml-quant.com/papers/arxiv/2412.12096/) (2024-12-18): The paper introduces PanSplat, a feed-forward method for wide-baseline panorama view synthesis supporting up to 4K resolution, featuring a unique spherical 3D Gaussian pyramid with a Fibonacci lattice arrangement for improved image quality and reduced information redundancy.
- [Apollo: An Exploration of Video Understanding in Large Multimodal Models](https://www.ml-quant.com/papers/arxiv/2412.10360/) (2024-12-18): The study explores the mechanisms behind video understanding in Large Multimodal Models (LMMs), showing that smaller models' design and training decisions effectively transfer to larger models, and introduces Apollo, a family of LMMs that excel across different model sizes.
- [The State of Robot Motion Generation](https://www.ml-quant.com/papers/arxiv/2410.12172/) (2024-12-18): The paper reviews 50 years of robotics research, discussing the evolution of methods for generating robot motion and the potential for integrating different techniques.
- [Flowedit: Inversion-Free Text-Based Editing Using Pre-Trained Flow Models](https://www.ml-quant.com/papers/arxiv/2412.08629/) (2024-12-18): Text-Based Editing: The study presents FlowEdit, a new text-based editing method for pre-trained text-to-image models, which outperforms the inversion approach by offering superior results at a lower transport cost.
- [ObjectMate: A Recurrence Prior for Object Insertion and Subject-Driven Generation](https://www.ml-quant.com/papers/arxiv/2412.08645/) (2024-12-18): Object Insertion Recurrence: The article introduces ObjectMate, a method for creating photorealistic compositions without altering the object's identity, eliminating the need for tuning.
- [Length Optimization in Conformal Prediction](https://www.ml-quant.com/papers/arxiv/2406.18814/) (2024-12-18): The research presents Conformal Prediction with Length-Optimization (CPL), a framework that creates optimal prediction sets while maintaining conditional validity under different covariate shifts.
- [Concept Bottleneck Language Models For protein design](https://www.ml-quant.com/papers/arxiv/2411.06090/) (2024-12-18): The paper presents Concept Bottleneck Protein Language Models (CB-pLM), a generative language model that provides control and interpretability in protein generation tasks without affecting performance.
- [Reinforcement Learning: An Overview](https://www.ml-quant.com/papers/arxiv/2412.05265/) (2024-12-12): The manuscript offers a detailed review of deep reinforcement learning and sequential decision making, covering value-based RL, policy-gradient methods, and model-based methods.
- [FlashRNN: I/O-Aware Optimization of Traditional RNNs on modern hardware](https://www.ml-quant.com/papers/arxiv/2412.07752/) (2024-12-12): The article introduces FlashRNN, a hardware-optimized solution for faster RNN processing, essential for time-series tasks and logical reasoning.
- [(MASK) is All You Need](https://www.ml-quant.com/papers/arxiv/2412.06787/) (2024-12-12): The study suggests using discrete-state models to connect Masked Generative and Non-autoregressive Diffusion models, and to redefine tasks like image segmentation as an unmasking process.
- [Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling](https://www.ml-quant.com/papers/arxiv/2412.05271/) (2024-12-12): The paper presents InternVL 2.5, a sophisticated multimodal large language model that performs well on various benchmarks, exceeding 70% on the MMMU benchmark.
- [Training Large Language Models to Reason in a Continuous Latent Space](https://www.ml-quant.com/papers/arxiv/2412.06769/) (2024-12-12): The article presents Coconut, a new approach that uses the last hidden state of large language models for reasoning in an unrestricted latent space, proving its effectiveness in enhancing the LLM on multiple reasoning tasks.
- [Efficient Online Reinforcement Learning Fine-Tuning Need Not Retain Offline Data](https://www.ml-quant.com/papers/arxiv/2412.07762/) (2024-12-12): The article introduces Warm-start RL (WSRL), a new reinforcement learning approach that doesn't require offline data, leading to quicker learning and better performance than previous algorithms.
- [Birth and Death of a Rose](https://www.ml-quant.com/papers/arxiv/2412.05278/) (2024-12-12): The research presents a technique for creating temporal object intrinsics like a blooming rose from pre-existing 2D diffusion models, allowing for the depiction of dynamic objects from any angle and lighting.
- [Chemist-aligned retrosynthesis by ensembling diverse inductive bias models.](https://www.ml-quant.com/papers/arxiv/2412.05269/) (2024-12-12): The paper suggests Chimera, a system for creating highly precise reaction models for chemical syntheses, which outperforms all major models by combining predictions from various sources using a learning-based ensembling strategy.
- [Extrapolated Urban View Synthesis Benchmark](https://www.ml-quant.com/papers/arxiv/2412.05256/) (2024-12-12): The study introduces the first Extrapolated Urban View Synthesis (EUVS) benchmark for assessing the performance of photorealistic simulators for self-driving vehicles, emphasizing the need for more robust methods and large-scale training.
- [P3-PO: Prescriptive Point Priors for Visuo-Spatial Generalization of Robot Policies](https://www.ml-quant.com/papers/arxiv/2412.06784/) (2024-12-12): The article presents Prescriptive Point Priors for Policies (P3-PO), a new framework that creates a unique state representation of the environment to enhance out-of-distribution generalization for robot manipulation, showing significant improvement over previous methods.
- [NVILA: Efficient Frontier Visual Language Models](https://www.ml-quant.com/papers/arxiv/2412.04468/) (2024-12-12): The article introduces NVILA, a new family of Visual Language Models (VLMs) that balances efficiency and accuracy, reducing training costs and latency while maintaining or improving accuracy.
- [Densing law of LLMs](https://www.ml-quant.com/papers/arxiv/2412.04315/) (2024-12-12): The article presents 'capacity density' as a new metric for evaluating Large Language Models (LLMs), showing that LLMs' capacity density doubles approximately every three months.
- [MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos](https://www.ml-quant.com/papers/arxiv/2412.04463/) (2024-12-12): The article discusses a deep visual SLAM framework that enables accurate, quick, and robust estimation of camera parameters and depth maps from casual monocular videos.
- [Monocular Dynamic Gaussian Splatting: Fast, Brittle, and Scene Complexity Rules](https://www.ml-quant.com/papers/arxiv/2412.04457/) (2024-12-12): The article provides an organized benchmark and analysis of Gaussian-splatting-based methods for converting multi-view image data into scene representations, offering comparisons not previously available.
- [Right on Time: Revising Time Series Models by Constraining their Explanations](https://www.ml-quant.com/papers/arxiv/2402.12921/) (2024-12-12): The article introduces Right on Time (RioT), a method that helps correct confounders in time series models by interacting with model explanations across both the time and frequency domain.
- [AlphaTablets: A Generic Plane Representation for 3D Planar Reconstruction from Monocular Videos](https://www.ml-quant.com/papers/arxiv/2411.19950/) (2024-12-04): 3D Reconstruction from Videos: AlphaTablets is a new 3D plane representation that merges the advantages of 2D and 3D models, providing accurate 3D plane modeling and superior performance in 3D planar reconstruction.
- [Critical Tokens Matter: Token-Level Contrastive Estimation Enhances LLM's Reasoning Capability](https://www.ml-quant.com/papers/arxiv/2411.19943/) (2024-12-04): Enhancing Reasoning: The cDPO method identifies and rewards 'critical tokens' that cause incorrect reasoning in Large Language Models, showing effectiveness in two popular models.
- [Another look at statistical inference with machine learning-imputed data](https://www.ml-quant.com/papers/arxiv/2411.19908/) (2024-12-04): The Chen and Chen estimator balances robustness and statistical efficiency in machine learning models, making it the top choice for prediction-based inference.
- [Incremental Multi-Scene Modeling via Continual Neural Graphics Primitives](https://www.ml-quant.com/papers/arxiv/2411.19903/) (2024-12-04): Modeling Scenes: The C^3-NeRF framework can incorporate multiple 3D scenes into a single neural radiance field, showing the ability to adapt to new scenes without needing old data or extra parameters.
- [LUMIA: Linear probing for Unimodal and MultiModal Membership Inference Attacks leveraging internal LLM states](https://www.ml-quant.com/papers/arxiv/2411.19876/) (2024-12-04): Membership Inference Attacks: LUMIA, a new method, uses Linear Probes to detect Membership Inference Attacks in Large Language Models, showing significant improvements over previous methods and providing insights into where attacks are most detectable.
- [Sparrow: Data-Efficient Video-LLM With Text-to-Image Augmentation](https://www.ml-quant.com/papers/arxiv/2411.19951/) (2024-12-04): The T2Vid method, developed by researchers, uses pre-trained image-LLMs to enhance video understanding, performing as well or better than full video datasets with only 15% of the sample size.
- [FreeCloth: Free-form Generation Enhances Challenging Clothed Human Modeling](https://www.ml-quant.com/papers/arxiv/2411.19942/) (2024-12-04): A new hybrid framework has been proposed for modeling clothed humans, using different strategies for various body regions, resulting in superior visual fidelity and realism.
- [A Graph-Based Classical and Quantum Approach to Deterministic L-System Inference](https://www.ml-quant.com/papers/arxiv/2411.19906/) (2024-12-04): A method for deducing deterministic context-free L-systems from a string sequence has been introduced, providing both a classical exact algorithm and an approximate quantum algorithm.
- [Adaptive Informed Deep Neural Networks for Power Flow Analysis](https://www.ml-quant.com/papers/arxiv/2412.02659/) (2024-12-04): The study presents PINN4PF, a deep learning structure for power flow analysis that effectively captures the nonlinear dynamics of large-scale modern power systems, surpassing both linear regression models and black-box NN.
- [On Domain-Adaptive Post-Training for Multimodal Large Language Models](https://www.ml-quant.com/papers/arxiv/2411.19930/) (2024-12-04): The paper explores domain adaptation of multimodal large language models through post-training, focusing on data synthesis, training pipelines, and task evaluation, resulting in improved domain-specific performance.
- [The Limits of Inference Scaling Through Resampling](https://www.ml-quant.com/papers/arxiv/2411.17501/) (2024-12-04): The study suggests that the accuracy of weaker language models cannot be indefinitely improved through inference scaling due to an unavoidable probability of false positives.
- [On the consistency of hyper-parameter selection in value-based deep reinforcement learning](https://www.ml-quant.com/papers/arxiv/2406.17523/) (2024-12-04): The paper explores the reliability of hyper-parameter selection in value-based deep reinforcement learning agents, introducing a new score to measure the consistency and reliability of different hyper-parameters.
- [Two Tales of Single-Phase Contrastive Hebbian Learning](https://www.ml-quant.com/papers/arxiv/2402.08573/) (2024-12-04): The authors lay the groundwork for the dual propagation method, a local learning algorithm for artificial neurons, and highlight its stability in relation to a specific adjoint state method, regardless of asymmetric nudging.
- [MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion Scaffolds](https://www.ml-quant.com/papers/arxiv/2405.17421/) (2024-12-04): The 4D Motion Scaffolds (MoSca) system uses vision models to create new views of dynamic scenes from single-view videos.
- [LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality Representation](https://www.ml-quant.com/papers/arxiv/2411.04997/) (2024-12-04): Enhancing Visual Representation: Large language models are integrated with the pretrained CLIP visual encoder, enhancing its ability to process complex captions.
- [Textured Gaussians for Enhanced 3D Scene Appearance Modeling](https://www.ml-quant.com/papers/arxiv/2411.18625/) (2024-12-04): A new Gaussian appearance representation is proposed, using texture maps to improve the expressivity of 3D Gaussian Splatting.
- [Diffusion Self-Distillation for Zero-Shot Customized Image Generation](https://www.ml-quant.com/papers/arxiv/2411.18616/) (2024-12-04): Diffusion Self-Distillation uses a pre-trained model to generate its own dataset for text-conditioned image-to-image tasks, offering more control for artists.
- [Video-Guided Foley Sound Generation with Multimodal Controls](https://www.ml-quant.com/papers/arxiv/2411.17698/) (2024-12-04): Video-Guided Sound Generation: MultiFoley, a new model for video-guided sound generation, allows users to create a variety of sound effects for videos using text, audio, and video inputs.
- [Multimodal Autoregressive Pre-training of Large Vision Encoders](https://www.ml-quant.com/papers/arxiv/2411.14402/) (2024-11-27): The article presents AIMV2, a method for training large-scale vision encoders using both images and text, which excels in multimodal image understanding and vision benchmarks.
- [OminiControl: Minimal and Universal Control for Diffusion Transformer](https://www.ml-quant.com/papers/arxiv/2411.15098/) (2024-11-27): OminiControl, a new framework that incorporates image conditions into pre-trained Diffusion Transformer models, is introduced, surpassing existing models in conditional generation tasks.
- [Marco-o1: Towards Open Reasoning Models for Open-Ended Solutions](https://www.ml-quant.com/papers/arxiv/2411.14405/) (2024-11-27): The research investigates the OpenAI o1 model's ability, enhanced by Chain-of-Thought fine-tuning and innovative reasoning strategies, to adapt to wider domains lacking clear standards.
- [Learning Humanoid Locomotion with Perceptive Internal Model](https://www.ml-quant.com/papers/arxiv/2411.14386/) (2024-11-27): The article introduces the Perceptive Internal Model (PIM), a method that uses elevation maps for stable humanoid robot movement across different terrains and sensor setups.
- [Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models](https://www.ml-quant.com/papers/arxiv/2411.14432/) (2024-11-27): The paper introduces Insight-V, a system that improves the reasoning abilities of large language models by generating extensive reasoning paths and integrating a multi-agent system for better visual reasoning performance.
- [XGrammar: Flexible and Efficient Structured Generation Engine for Large Language Models](https://www.ml-quant.com/papers/arxiv/2411.15100/) (2024-11-27): The article introduces XGrammar, a structure generation engine that significantly speeds up context-free grammar execution in large language models.
- [Whack-a-Chip: The Futility of Hardware-Centric Export Controls](https://www.ml-quant.com/papers/arxiv/2411.14425/) (2024-11-27): The study reveals how Chinese companies like Tencent are bypassing U.S. export controls to use semiconductors in advanced AI models.
- [Material Anything: Generating Materials for Any 3D Object via Diffusion](https://www.ml-quant.com/papers/arxiv/2411.15138/) (2024-11-27): Material Anything is a new framework that creates physically-based materials for 3D objects, adaptable to various lighting conditions.
- [Stable Flow: Vital Layers for Training-Free Image Editing](https://www.ml-quant.com/papers/arxiv/2411.14430/) (2024-11-27): The research introduces a method to identify crucial layers in Diffusion Transformer models, improving image editing and inversion methods.
- [Exploring Discrete Flow Matching for 3D De Novo Molecule Generation](https://www.ml-quant.com/papers/arxiv/2411.16644/) (2024-11-27): The study presents FlowMol-CTMC, a model that outperforms others in 3D small molecule generation with fewer learnable parameters.
- [REDUCIO! Generating 1K Video Within 16 Seconds Using Extremely Compressed Motion Latents](https://www.ml-quant.com/papers/arxiv/2411.13552/) (2024-11-27): The article discusses Reducio-DiT, a technique that encodes videos into a compressed format, enabling the creation of high-resolution videos with limited GPU resources.
- [Retrieval with Learned Similarities](https://www.ml-quant.com/papers/arxiv/2407.15462/) (2024-11-27): The paper introduces Mixture-of-Logits (MoL), a method that enhances the performance of recommendation systems and language models by approximating similarity functions, reducing latency by up to 66 times.
- [KTO: Model Alignment as Prospect Theoretic Optimization](https://www.ml-quant.com/papers/arxiv/2402.01306/) (2024-11-27): The study presents KTO, a new method that uses a Kahneman-Tversky model to align language models with human feedback, performing better than preference-based methods by learning from a binary signal of output desirability.
- [Disentangling Memory and Reasoning Ability in Large Language Models](https://www.ml-quant.com/papers/arxiv/2411.13504/) (2024-11-27): The article presents a new approach for Large Language Models that divides the process into memory recall and reasoning, enhancing model performance and interpretability.
- [Basic syntax from speech: Spontaneous concatenation in unsupervised deep neural networks](https://www.ml-quant.com/papers/arxiv/2305.01626/) (2024-11-27): The paper discusses spontaneous concatenation in convolutional neural networks trained on acoustic recordings, offering a potential neural method for modeling syntax from raw acoustic inputs.
- [MaGS: Reconstructing and Simulating Dynamic 3D Objects with Mesh-Adsorbed Gaussian Splatting](https://www.ml-quant.com/papers/arxiv/2406.01593/) (2024-11-27): The research presents the Mesh-adsorbed Gaussian Splatting method for 3D reconstruction, combining 3D Gaussians and meshes for improved performance.
- [MagicQuill: An Intelligent Interactive Image Editing System](https://www.ml-quant.com/papers/arxiv/2411.09703/) (2024-11-20): Image Editing System: MagicQuill is an image editing system that uses a large language model to predict editing intentions in real time, enabling quick and accurate image modifications.
- [NeuralDEM - Real-time Simulation of Industrial Particulate Flows](https://www.ml-quant.com/papers/arxiv/2411.09678/) (2024-11-20): Particulate Flow Simulation: NeuralDEM is a deep learning approach that replaces slow routines in the discrete element method (DEM), allowing for quicker and more efficient simulations of large fluid-mechanical and particulate systems.
- [LlaVA-CoT: Let Vision Language Models Reason Step-By-Step](https://www.ml-quant.com/papers/arxiv/2411.10440/) (2024-11-20): Vision Language Model: LLaVA-o1 is a new Vision-Language Model that performs autonomous multistage reasoning, enhancing precision in reasoning-intensive tasks and surpassing larger models in various multimodal reasoning benchmarks.
- [Squeezed Attention: Accelerating Long Context Length LLM Inference](https://www.ml-quant.com/papers/arxiv/2411.09688/) (2024-11-20): LLM Inference: Squeezed Attention is a proposed method to speed up Large Language Model applications by using K-means clustering to group similar keys in fixed context inputs, reducing computational costs and enhancing inference efficiency.
- [Adaptive Decoding via Latent Preference Optimization](https://www.ml-quant.com/papers/arxiv/2411.09661/) (2024-11-20): Preference Optimization: Adaptive Decoding is a technique that dynamically selects the sampling temperature during language model decoding, optimizing performance across various tasks that require different temperatures.
- [Probing LLM Hallucination from Within: Perturbation-Driven Approach via Internal Knowledge](https://www.ml-quant.com/papers/arxiv/2411.09689/) (2024-11-20): A new task called Hallucination Reasoning is introduced to better categorize text generated by Language Learning Models, enhancing the detection of unfaithful text generation.
- [How Do Machine Learning Models Change?](https://www.ml-quant.com/papers/arxiv/2411.09645/) (2024-11-20): A large-scale study of Hugging Face models reveals patterns in commit and release activities, highlighting the continuous improvement in machine learning models.
- [Towards a Classification of Open-Source ML Models and Datasets for Software Engineering](https://www.ml-quant.com/papers/arxiv/2411.09683/) (2024-11-20): A study classifies Pre-Trained Models and datasets on Hugging Face using a Software Engineering approach, indicating a need for more task coverage to better integrate machine learning in software engineering.
- [Image Matching Filtering and Refinement by Planes and Beyond](https://www.ml-quant.com/papers/arxiv/2411.09484/) (2024-11-20): A new non-deep learning method for image matching is introduced, showing superior or equivalent performance to recent state-of-the-art deep learning methods.
- [On the Foundation Model for Cardiac MRI Reconstruction](https://www.ml-quant.com/papers/arxiv/2411.10403/) (2024-11-20): A new foundation model for cardiac magnetic resonance imaging is proposed, which improves image quality across various protocols and outperforms traditional machine learning methods.
- [A Single Transformer for Scalable Vision-Language Modeling](https://www.ml-quant.com/papers/arxiv/2407.06438/) (2024-11-20): SOLO is a unified transformer for vision-language modeling, addressing scalability issues in large models and providing an open-source training blueprint.
- [Learning Diffusion Priors from Observations by Expectation Maximization](https://www.ml-quant.com/papers/arxiv/2405.13712/) (2024-11-20): A new method using the expectation-maximization algorithm has been developed to train diffusion models from incomplete and noisy data, enhancing their effectiveness for subsequent tasks.
- [Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data](https://www.ml-quant.com/papers/arxiv/2404.03862/) (2024-11-20): Quote-Tuning is a new approach that prompts large language models to directly quote from reliable sources, thereby improving their credibility and verifiability.
- [Vector Retrieval with Similarity and Diversity: How Hard Is It?](https://www.ml-quant.com/papers/arxiv/2407.04573/) (2024-11-20): The article presents a new method for vector retrieval in Large Language Models (LLMs) called VRSD, which ensures similarity and diversity constraints and performs better than the commonly used MMR method.
- [4D Gaussian Splatting in the Wild with Uncertainty-Aware Regularization](https://www.ml-quant.com/papers/arxiv/2411.08879/) (2024-11-20): The authors introduce a 4D Gaussian Splatting (4DGS) algorithm that enhances the synthesis of novel views from casually recorded monocular videos, improving image reconstruction quality.
- [Dynamic Rewarding with Prompt Optimization Enables Tuning-free Self-Alignment of Language Models](https://www.ml-quant.com/papers/arxiv/2411.08733/) (2024-11-20): The paper presents a new tuning-free method called Dynamic Rewarding with Prompt Optimization (DRPO) for self-aligning Large Language Models (LLMs), improving alignment performance without extra training or human intervention.
- [Large Language Model-Based Interpretable Machine Learning Control in Building Energy Systems](https://www.ml-quant.com/papers/arxiv/2402.09584/) (2024-11-20): The study investigates the use of Interpretable Machine Learning (IML) in HVAC systems to enhance transparency and understanding, using Shapley values and Large Language Models (LLMs) to create a comprehensible narrative.
- [The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities](https://www.ml-quant.com/papers/arxiv/2411.04986/) (2024-11-13): Modern language models can process various languages and forms by learning a shared representation space used during input processing.
- [Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models](https://www.ml-quant.com/papers/arxiv/2411.04996/) (2024-11-13): The Mixture-of-Transformers (MoT) is a sparse multi-modal transformer architecture that reduces pretraining costs and allows modality-specific processing with global self-attention.
- [Enabling LLM Knowledge Analysis via Extensive Materialization](https://www.ml-quant.com/papers/arxiv/2411.04920/) (2024-11-13): A large general-domain knowledge base (GPTKB) can be built entirely from a large language model, containing 105 million triples for over 2.9 million entities.
- [Non-equilibrium active noise enhances generative memory in diffusion models](https://www.ml-quant.com/papers/arxiv/2411.07233/) (2024-11-13): The generative performance of diffusion models can be enhanced by using noise sources with temporal correlations for data destruction in the forward process.
- [ReCapture: Generative Video Camera Controls for User-Provided Videos using Masked Video Fine-Tuning](https://www.ml-quant.com/papers/arxiv/2411.05003/) (2024-11-13): ReCapture is a method for creating new videos with unique camera trajectories from a single video, allowing for the regeneration of the video from different angles and cinematic camera motion.
- [Dynamem: Online Dynamic Spatio-Semantic Memory for Open World Mobile Manipulation](https://www.ml-quant.com/papers/arxiv/2411.04999/) (2024-11-13): Mobile Manipulation: DynaMem is a new method that uses dynamic spatio-semantic memory to help robots explore and adapt to new environments and track object movements.
- [Stronger Models are NOT Stronger Teachers for Instruction Tuning](https://www.ml-quant.com/papers/arxiv/2411.07133/) (2024-11-13): A study introduces Compatibility-Adjusted Reward (CAR), a new metric to evaluate the effectiveness of language models, challenging the belief that larger models are better for instruction tuning.
- [Watermark Anything with Localized Messages](https://www.ml-quant.com/papers/arxiv/2411.07231/) (2024-11-13): The Watermark Anything Model (WAM) is a deep-learning model that can embed and extract hidden watermarks in specific areas of an image.
- [RefreshKV: Updating Small KV Cache During Long-form Generation](https://www.ml-quant.com/papers/arxiv/2411.05787/) (2024-11-13): Recycled Attention is a method for large language models that alternates attention between full context and a subset of input tokens, improving performance and reducing computational load in long-context tasks.
- [Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models](https://www.ml-quant.com/papers/arxiv/2411.07232/) (2024-11-13): Object Insertion: Add-it is a training-free approach for semantic image editing that uses diffusion models to add objects into images based on text instructions, ensuring natural placement and detail preservation.
- [Logits of API-Protected LLMs Leak Proprietary Information](https://www.ml-quant.com/papers/arxiv/2403.09539/) (2024-11-13): Researchers have discovered a method to extract hidden information from large language models like OpenAI's gpt-3.5-turbo using API queries, exploiting a weakness known as the softmax bottleneck.
- [Is ChatGPT Transforming Academics' Writing Style?](https://www.ml-quant.com/papers/arxiv/2404.08627/) (2024-11-13): A study of a million arXiv papers shows that large language models, specifically ChatGPT, are significantly impacting the writing style of academic abstracts, especially in computer science.
- [No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations](https://www.ml-quant.com/papers/arxiv/2407.10964/) (2024-11-13): The FUNGI method improves transformer encoders' features using self-supervised gradients, enhancing performance in vision, natural language processing, and audio tasks and datasets.
- [Nteasee: Understanding Needs in AI for Health in Africa - A Mixed-Methods Study of Expert and General Population Perspectives](https://www.ml-quant.com/papers/arxiv/2409.12197/) (2024-11-13): Research highlights the potential of AI in African healthcare, but emphasizes the need for culturally sensitive approaches and addresses concerns about trust, ethics, and systemic barriers.
- [Optimization without Retraction on the Random Generalized Stiefel Manifold](https://www.ml-quant.com/papers/arxiv/2405.01702/) (2024-11-13): A new method for optimization over matrices is proposed, which is cost-effective and efficient in various machine learning applications.
- [A Implies B: Circuit Analysis in LLMs for Propositional Logical Reasoning](https://www.ml-quant.com/papers/arxiv/2411.04105/) (2024-11-13): A study reveals the internal mechanisms of large language models that enable complex logical reasoning, identifying specific planning and reasoning circuits.
- [Meta-models for transfer learning in source localization](https://www.ml-quant.com/papers/arxiv/2305.08657/) (2024-11-13): A Bayesian multilevel approach is used to predict model hyperparameters in acoustic emission experiments, demonstrating its use in source localization.
- [SynCode: LLM Generation with Grammar Augmentation](https://www.ml-quant.com/papers/arxiv/2403.01632/) (2024-11-13): LLM Generation with Grammar Augmentation: SynCode, a new framework for syntactical decoding with large language models, is introduced, significantly reducing syntax errors in Python and Go code generation.
- ["Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization](https://www.ml-quant.com/papers/arxiv/2411.02355/) (2024-11-06): A study shows that FP8 weight and activation quantization in large language models (LLMs) is lossless across all scales, while INT8 and INT4 quantizations also perform well with proper tuning, offering guidelines for deploying quantized LLMs.
- [URAvatar: Universal Relightable Gaussian Codec Avatars](https://www.ml-quant.com/papers/arxiv/2410.24223/) (2024-11-06): A novel method creates photorealistic, relightable head avatars from a phone scan with unknown illumination, surpassing existing methods and maintaining real-time rendering.
- [SelfCodeAlign: Self-Alignment for Code Generation](https://www.ml-quant.com/papers/arxiv/2410.24198/) (2024-11-06): SelfCodeAlign, a new pipeline, enhances the ability of LLMs to follow human instructions without extensive human annotations, outperforming previous methods and creating a top-performing coding model.
- [EgoMimic: Scaling Imitation Learning via Egocentric Video](https://www.ml-quant.com/papers/arxiv/2410.24221/) (2024-11-06): EgoMimic, a new framework, improves manipulation tasks performance using human embodiment data, proving more effective than existing imitation learning methods and showing that human data is more valuable than robot data.
- [GeoSplatting: Towards Geometry Guided Gaussian Splatting for Physically-Based Inverse Rendering](https://www.ml-quant.com/papers/arxiv/2410.24204/) (2024-11-06): GeoSplatting, a new hybrid representation combining 3D Gaussian Splatting with geometric guidance and differentiable physically-based rendering equations, enhances the accuracy of physically-based inverse rendering, outperforming existing methods.
- [Towards Generative Ray Path Sampling for Faster Point-to-Point Ray Tracing](https://www.ml-quant.com/papers/arxiv/2410.23773/) (2024-11-06): The article introduces a Machine Learning-enhanced Ray Tracing method for efficient radio propagation modeling, reducing computational effort while maintaining accuracy.
- [Machine learning identification of maternal inflammatory response and histologic choroamnionitis from placental membrane whole slide images](https://www.ml-quant.com/papers/arxiv/2411.02354/) (2024-11-06): The paper explores the use of machine learning to analyze Maternal Inflammatory Response (MIR) from whole slide images, achieving up to 88.5% balanced accuracy.
- [Oblivious Defense in ML Models: Backdoor Removal without Detection](https://www.ml-quant.com/papers/arxiv/2411.03279/) (2024-11-06): The research proposes strategies to defend against undetectable backdoors in machine learning models, using random self-reducibility-inspired techniques.
- [Attacking Vision-Language Computer Agents via Pop-ups](https://www.ml-quant.com/papers/arxiv/2411.02391/) (2024-11-06): The study shows that autonomous agents using large vision and language models can be significantly disrupted by strategically designed adversarial pop-ups.
- [How Far is Video Generation from World Model: A Physical Law Perspective](https://www.ml-quant.com/papers/arxiv/2411.02385/) (2024-11-06): A Physical Law Perspective: The paper assesses the capability of video generation models to identify physical laws from visual data, indicating that mere scaling is not enough for these models to discover fundamental physical laws.
- [Efficient Adversarial Training in LLMs with Continuous Attacks](https://www.ml-quant.com/papers/arxiv/2405.15589/) (2024-11-06): CAdvUL, a new adversarial training algorithm, enhances the resilience of large language models against adversarial attacks by efficiently calculating attacks in the continuous embedding space.
- [NAVSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Benchmarking](https://www.ml-quant.com/papers/arxiv/2406.15349/) (2024-11-06): Autonomous Vehicle Simulation: NAVSIM, a non-reactive simulator, allows for large-scale real-world testing of vision-based driving policies using extensive datasets and simulation-based metrics, bridging the gap between open-loop and closed-loop evaluations.
- [AmbigNLG: Addressing Task Ambiguity in Instruction for NLG](https://www.ml-quant.com/papers/arxiv/2402.17717/) (2024-11-06): Task Ambiguity in NLG: AmbigNLG, a new task and dataset, tackles task ambiguity in instructions for Natural Language Generation, improving the alignment of generated text with user expectations and boosting the performance of Large Language Models.
- [HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale](https://www.ml-quant.com/papers/arxiv/2409.16299/) (2024-11-06): Coding Task SE Agents: HyperAgent, a multi-agent system for software engineering tasks, has achieved new benchmarks in tasks like GitHub issue resolution and code generation across different programming languages.
- [DuQuant: Distributing Outliers via Dual Transformation Makes Stronger Quantized LLMs](https://www.ml-quant.com/papers/arxiv/2406.01721/) (2024-11-06): Outlier Management for LLMs: DuQuant, a new method for quantizing large language models, uses rotation and permutation transformations to manage outliers, surpassing performance of existing baselines in various tasks.
- [LaCour!: enabling research on argumentation in hearings of the European Court of Human Rights](https://www.ml-quant.com/papers/arxiv/2312.05061/) (2024-11-06): Argumentation Research in ECHR Hearings: LaCour!, the first corpus of textual oral arguments from the European Court of Human Rights, provides transcribed multilingual oral hearings linked to final judgments, enhancing legal research.
- [Expert-aided causal discovery of ancestral graphs](https://www.ml-quant.com/papers/arxiv/2309.12032/) (2024-11-06): A novel method for causal inference uses ancestral graph sampling and expert feedback to refine causal discovery, providing uncertainty estimates and accounting for unobserved confounders.
- [On the Benefits of Active Data Collection in Operator Learning](https://www.ml-quant.com/papers/arxiv/2410.19725/) (2024-10-31): The study shows that active data collection methods are more effective than passive ones in operator learning, especially when dealing with linear operators and input functions from a mean-zero stochastic process.
- [Proportional Fairness in Non-Centroid Clustering](https://www.ml-quant.com/papers/arxiv/2410.23273/) (2024-10-31): The research introduces a new algorithm, GreedyCohesiveClustering, that expands fair clustering to non-centroid clustering, and an auditing algorithm to measure fairness approximation.
- [Understanding Synthetic Context Extension via Retrieval Heads](https://www.ml-quant.com/papers/arxiv/2410.22316/) (2024-10-31): The paper finds that fine-tuning long-context language models with synthetic data improves performance in retrieval and reasoning tasks, with attention heads predicting performance.
- [Multi-modal AI for comprehensive breast cancer prognostication](https://www.ml-quant.com/papers/arxiv/2410.21256/) (2024-10-31): A new AI test for breast cancer patient stratification, combining digital pathology and clinical characteristics, has been developed, showing higher accuracy than the current standard and applicability across all major breast cancer subtypes.
- [Batch, match, and patch: low-rank approximations for score-based variational inference](https://www.ml-quant.com/papers/arxiv/2410.22292/) (2024-10-31): The paper introduces an extension of the batch-and-match framework for black-box variational inference to high-dimensional problems, using compact parameterization of full covariance matrices to enhance efficiency and performance.
- [An Efficient Approach to Generate Safe Drivable Space by LiDAR-Camera-HDmap Fusion](https://www.ml-quant.com/papers/arxiv/2410.22314/) (2024-10-31): The article suggests a perception module for self-driving cars using LiDAR, camera, and HD map data fusion for accurate drivable space detection in all weather conditions.
- [FISHNET: Financial Intelligence from Sub-querying, Harmonizing, Neural-Conditioning, Expert Swarms, and Task Planning](https://www.ml-quant.com/papers/doi/10-1145-3677052-3698597/) (2024-10-31): Financial Intelligence: The paper introduces FISHNET, a new system for generating financial intelligence from large data sources, offering scalability, flexibility, and data integrity.
- [Conditional Forecasting of Margin Calls using Dynamic Graph Neural Networks](https://www.ml-quant.com/papers/arxiv/2410.23275/) (2024-10-31): The authors propose a Dynamic Graph Neural Network for predicting in temporal financial networks, offering a tool for systemic risk monitoring in financial entities trading swap contracts.
- [Arabic Music Classification and Generation using Deep Learning](https://www.ml-quant.com/papers/arxiv/2410.19719/) (2024-10-31): The study suggests a machine learning method using a convolutional neural network for classifying and creating new and traditional Egyptian music by composer, with 81.4% accuracy.
- [Modular Duality in Deep Learning](https://www.ml-quant.com/papers/arxiv/2410.21265/) (2024-10-31): The article presents a new theory of modular dualization for general neural networks, providing a theoretical basis for fast and scalable training algorithms, potentially leading to a new generation of optimizers for neural architectures.
- [Empirical Design in Reinforcement Learning](https://www.ml-quant.com/papers/arxiv/2304.01315/) (2024-10-31): The article highlights the importance of proper statistical evidence and avoiding common errors in empirical design for effective reinforcement learning experiments.
- [Unbounded: A Generative Infinite Game of Character Life Simulation](https://www.ml-quant.com/papers/arxiv/2410.18975/) (2024-10-31): A Generative Game: The paper presents Unbounded, a generative infinite game using a large language model and a dynamic image prompt Adapter for real-time game creation.
- [Emergent mechanisms for long timescales depend on training curriculum and affect performance in memory tasks](https://www.ml-quant.com/papers/arxiv/2309.12927/) (2024-10-31): The research shows that recurrent neural networks improve performance and generalization by adapting timescales for memory-dependent tasks.
- [TabReD: Analyzing Pitfalls and Filling the Gaps in Tabular Deep Learning Benchmarks](https://www.ml-quant.com/papers/arxiv/2406.19380/) (2024-10-31): The article presents TabReD, a collection of industry-grade tabular datasets, showing that simple MLP-like architectures and GBDT perform best in real-world conditions.
- [Adam with model exponential moving average is effective for nonconvex optimization](https://www.ml-quant.com/papers/arxiv/2405.18199/) (2024-10-31): The article analyzes the optimal convergence rates of two optimization techniques, Adam and model exponential moving average (EMA), in different nonconvex optimization settings.
- [Endoscapes, a critical view of safety and surgical scene segmentation dataset for laparoscopic cholecystectomy](https://www.ml-quant.com/papers/arxiv/2312.12429/) (2024-10-31): The report discusses Endoscapes, a dataset of annotated laparoscopic cholecystectomy videos, designed for automated assessment of the Critical View of Safety.
- [Block and Detail: Scaffolding Sketch-to-Image Generation](https://www.ml-quant.com/papers/arxiv/2402.18116/) (2024-10-31): The paper presents a sketch-to-image tool that can generate high-quality images from sketches, showcasing its effectiveness through various examples and comparisons.
- [PixelGaussian: Generalizable 3D Gaussian Reconstruction from Arbitrary Views](https://www.ml-quant.com/papers/arxiv/2410.18979/) (2024-10-31): The authors introduce PixelGaussian, a framework that learns 3D Gaussian reconstruction from any view, adjusting the Gaussian distribution and quantity based on geometric complexity.
- [DisC-GS: Discontinuity-aware Gaussian Splatting](https://www.ml-quant.com/papers/arxiv/2405.15196/) (2024-10-31): The paper presents a new framework that allows Gaussian Splatting to accurately render discontinuities and boundaries in images, addressing its previous limitations.
- [Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens](https://www.ml-quant.com/papers/arxiv/2410.13863/) (2024-10-23): The study explores text-to-image generation, finding continuous token-based models offer superior visual quality and random-order models score higher on the GenEval benchmark, leading to a new model, Fluid.
- [Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation](https://www.ml-quant.com/papers/arxiv/2410.13848/) (2024-10-23): The paper presents Janus, a framework that separates visual encoding into different pathways for multimodal understanding and generation, offering improved performance and flexibility.
- [Bridging the Training-Inference Gap in LLMs by Leveraging Self-Generated Tokens](https://www.ml-quant.com/papers/arxiv/2410.14655/) (2024-10-23): The paper suggests two methods to address the discrepancy between training and inference time in language models, resulting in enhanced performance in tasks like summarization and question-answering.
- [DepthSplat: Connecting Gaussian Splatting and Depth](https://www.ml-quant.com/papers/arxiv/2410.13862/) (2024-10-23): The study introduces DepthSplat, a model that combines Gaussian splatting and depth estimation, leading to improved performance in depth estimation and novel view synthesis.
- [Differentiable Robot Rendering](https://www.ml-quant.com/papers/arxiv/2410.13851/) (2024-10-23): The paper presents a method for differentiable robot rendering, enabling the visual appearance of a robot to be directly differentiable with respect to its control parameters, useful for reconstructing robot poses from images and controlling robots through vision language models.
- [The Disparate Benefits of Deep Ensembles](https://www.ml-quant.com/papers/arxiv/2410.13831/) (2024-10-23): Deep Ensembles, a type of AI, can unintentionally favor certain groups, leading to unfair benefits; this can be reduced through post-processing without affecting performance.
- [xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs](https://www.ml-quant.com/papers/arxiv/2410.16267/) (2024-10-23): XGen-MM-Vid (BLIP-3-Video) is a language model for videos that captures temporal information efficiently, offering accuracy similar to larger models but with greater efficiency.
- [Knowledge Transfer from Simple to Complex: A Safe and Efficient Reinforcement Learning Framework for Autonomous Driving Decision-Making](https://www.ml-quant.com/papers/arxiv/2410.14468/) (2024-10-23): The Simple to Complex Collaborative Decision framework uses reinforcement learning to enhance safety and efficiency in autonomous vehicle decision-making, guided by a teacher model to avoid danger.
- [Agent-to-Sim: Learning Interactive Behavior Models from Casual Longitudinal Videos](https://www.ml-quant.com/papers/arxiv/2410.16259/) (2024-10-23): Agent-to-Sim (ATS) is a system that learns interactive behavior models of 3D agents from video collections, allowing transfer from real-life videos to a behavior simulator.
- [LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation](https://www.ml-quant.com/papers/arxiv/2410.13846/) (2024-10-23): SimLayerKV is a technique that minimizes memory usage in large language models by identifying and reducing cache in lazy layers, achieving significant cache compression with minimal performance loss.
- [EasyRec: Simple yet Effective Language Models for Recommendation](https://www.ml-quant.com/papers/arxiv/2408.08821/) (2024-10-23): Recommendation Language Models: The study presents EasyRec, a method that combines text-based semantic understanding with collaborative signals for recommender systems, showing improved performance in text-based zero-shot recommendation situations.
- [Depth Any Video with Scalable Synthetic Data](https://www.ml-quant.com/papers/arxiv/2410.10815/) (2024-10-17): The article presents Depth Any Video, a new model that uses synthetic data and video diffusion models to estimate video depth more accurately and consistently than previous models.
- [Your Mixture-of-Experts LLM Is Secretly an Embedding Model For Free](https://www.ml-quant.com/papers/arxiv/2410.10814/) (2024-10-17): The research shows that Mixture-of-Experts Large Language Models can be effective embedding models without finetuning, and suggests a combination of routing weights and hidden state for better performance.
- [CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos](https://www.ml-quant.com/papers/arxiv/2410.11831/) (2024-10-17): Simple Tracking: The paper introduces CoTracker3, a new tracking model that uses semi-supervised training to generate pseudo-labels from real videos, improving point tracking performance with less data.
- [Long-LRM: Long-Sequence Large Reconstruction Model for Wide-Coverage Gaussian Splats](https://www.ml-quant.com/papers/arxiv/2410.12781/) (2024-10-17): Efficient 3D Reconstruction: The article discusses Long-LRM, a 3D Gaussian reconstruction model that can efficiently reconstruct large scenes from long sequences of input images, surpassing previous feed-forward models.
- [Mix Data or Merge Models? Optimizing for Diverse Multi-Task Learning](https://www.ml-quant.com/papers/arxiv/2410.10801/) (2024-10-17): The research investigates model merging in a multilingual context for Large Language Models, finding that objective-based and language-based merging methods enhance performance and safety.
- [Towards Foundation Models for 3D Vision: How Close are We?](https://www.ml-quant.com/papers/arxiv/2410.10799/) (2024-10-17): A new 3D visual understanding benchmark shows that while specialized models are accurate, they are not robust, and human vision is still the most reliable 3D visual system.
- [4‐LEGS: 4D Language Embedded Gaussian Splatting](https://www.ml-quant.com/papers/arxiv/2410.10719/) (2024-10-17): 4D Language Embedded Gaussian Splatting: A new method uses 4D representation to connect language with a dynamic model of the world, enabling users to locate events in a video from text prompts.
- [Variance reduction combining pre-experiment and in-experiment data](https://www.ml-quant.com/papers/arxiv/2410.09027/) (2024-10-17): A new method combining pre-experiment and in-experiment data increases the sensitivity of online controlled experiments, improving variance reduction and speeding up decision-making.
- [Geometry-Aware Generative Autoencoders for Warped Riemannian Metric Learning and Generative Modeling on Data Manifolds](https://www.ml-quant.com/papers/arxiv/2410.12779/) (2024-10-17): The Geometry-Aware Generative Autoencoder (GAGA) addresses challenges of high-dimensional datasets by combining manifold learning with generative modeling.
- [Improving Long-Text Alignment for Text-to-Image Diffusion Models](https://www.ml-quant.com/papers/arxiv/2410.11817/) (2024-10-17): LongAlign, a new method for processing long texts, improves alignment in text-to-image diffusion models, overcoming limitations of existing encoding methods.
- [Scaling Laws For Diffusion Transformers](https://www.ml-quant.com/papers/arxiv/2410.08184/) (2024-10-17): Experiments have confirmed the existence of scaling laws in Diffusion Transformers, aiding in determining optimal model size, data needs, and predicting text-to-image generation loss.
- [Reuse Your Rewards: Reward Model Transfer for Zero-Shot Cross-Lingual Alignment](https://www.ml-quant.com/papers/arxiv/2404.12318/) (2024-10-17): A study has found that language models aligned with human-annotated preference data are preferred by humans in over 70% of cases, even without language-specific data for supervised finetuning.
- [Stability-Aware Training of Machine Learning Force Fields with Differentiable Boltzmann Estimators](https://www.ml-quant.com/papers/arxiv/2402.13984/) (2024-10-17): The Stability-Aware Boltzmann Estimator Training has been introduced to improve the stability, data efficiency, and agreement with reference observables in Machine Learning Force Fields used in molecular dynamics simulations.
- [An exactly solvable model for emergence and scaling laws in the multitask sparse parity problem](https://www.ml-quant.com/papers/arxiv/2404.17563/) (2024-10-17): A new framework represents each new ability in deep learning models as a basis function, providing analytic expressions for the emergence of new skills and scaling laws of the loss with various factors.
- [Evaluating Copyright Takedown Methods for Language Models](https://www.ml-quant.com/papers/arxiv/2406.18664/) (2024-10-17): The article discusses CoTaEval, a new framework for evaluating methods to prevent AI from generating copyrighted content, highlighting the need for further research as no method was found to be completely effective.
- [Shielded Diffusion: Generating Novel and Diverse Images using Sparse Repellency](https://www.ml-quant.com/papers/arxiv/2410.06025/) (2024-10-17): The paper introduces SPELL, a method that enhances the diversity of text-to-image diffusion models while avoiding protected images, proving its superiority over other diversity methods.
- [cedar: Optimized and Unified Machine Learning Input Data Pipelines](https://www.ml-quant.com/papers/arxiv/2401.08895/) (2024-10-17): The study introduces cedar, a new programming framework for machine learning input data pipelines that enhances performance by applying a mix of optimizations, outperforming other systems.
- [Towards Scalable Exact Machine Unlearning Using Parameter-Efficient Fine-Tuning](https://www.ml-quant.com/papers/arxiv/2406.16257/) (2024-10-17): The article presents Sequence-aware Sharded Sliced Training (S3T), a new framework for machine unlearning that improves system deletion capabilities with minimal impact on model performance, proving more effective than other methods.
- [Learning Quadruped Locomotion Using Differentiable Simulation](https://www.ml-quant.com/papers/arxiv/2403.14864/) (2024-10-17): The paper proposes a new differentiable simulation framework for learning quadruped locomotion, showing its efficiency and effectiveness compared to traditional reinforcement learning methods.
- [Differential Transformer](https://www.ml-quant.com/papers/arxiv/2410.05258/) (2024-10-09): The Diff Transformer is a new language model architecture that enhances attention to relevant context and reduces noise, outperforming the standard Transformer in long-context modeling and key information retrieval.
- [GS-VTON: Controllable 3D Virtual Try-on with Gaussian Splatting](https://www.ml-quant.com/papers/arxiv/2410.05259/) (2024-10-09): GS-VTON improves 3D virtual try-on by transferring knowledge from 2D models, ensuring consistency across different viewpoints and enhancing the quality of 3D geometry.
- [Bias-Aware Conformal Prediction for Metric-Based Imaging Pipelines](https://www.ml-quant.com/papers/arxiv/2410.05263/) (2024-10-09): A study on Conformal Prediction intervals shows that asymmetric adjustments are unaffected by bias and maintain the same validity as if the bias never occurred, unlike symmetric adjustments.
- [MA-RLHF: Reinforcement Learning from Human Feedback with Macro Actions](https://www.ml-quant.com/papers/arxiv/2410.02743/) (2024-10-09): The MA-RLHF framework integrates macro actions into the learning process of large language models, enhancing learning efficiency and performance in tasks like text summarization and dialogue generation.
- [GenSim2: Scaling Robot Data Generation with Multi-modal and Reasoning LLMs](https://www.ml-quant.com/papers/arxiv/2410.03645/) (2024-10-09): GenSim2 is a scalable framework for robotic simulation that uses large language models for task creation and a multi-task language-conditioned policy architecture to learn from demonstrations, improving policy performance and enabling zero-shot transfer.
- [GPT-4o as the Gold Standard: A Scalable and General Purpose Approach to Filter Language Model Pretraining Data](https://www.ml-quant.com/papers/arxiv/2410.02755/) (2024-10-09): Data Filtering System with GPT-4o Accuracy: The article introduces SIEVE, a cost-effective method for filtering web-scale data that matches the accuracy of GPT-4o and is efficient in curating large datasets for language model training.
- [Language Model Training on Edit Sequences Enhances Code Synthesis](https://www.ml-quant.com/papers/web/9b37a775db/) (2024-10-09): The paper presents LintSeq, a synthetic data generation algorithm that refactors code into a sequence of edits, resulting in more diverse programs and improved code synthesis performance.
- [TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens](https://www.ml-quant.com/papers/arxiv/2410.05261/) (2024-10-09): Efficient LVLM for Bilingual OCR: The study introduces TextHawk2, a bilingual Large Vision-Language Model that provides efficient fine-grained perception and superior performance with 16 times fewer image tokens than previous models.
- [Function-Guided Conditional Generation Using Protein Language Models with Adapters](https://www.ml-quant.com/papers/arxiv/2410.03634/) (2024-10-09): The research proposes ProCALM, a method for generating proteins conditionally using adapters to protein language models, capable of generating sequences from target enzyme families and generalizing to rare and unseen ones.
- [LoTLIP: Improving Language-Image Pre-training for Long Text Understanding](https://www.ml-quant.com/papers/arxiv/2410.05249/) (2024-10-09): Enhancing Language-Image Pre-training: The article reveals that language-image pre-training models struggle with long text due to training images being paired with short captions, and suggests a solution that improves long-text image retrieval by 11.1%.
- [Training Language Models to Self-Correct via Reinforcement Learning](https://www.ml-quant.com/papers/arxiv/2409.12917/) (2024-10-09): SCoRe, a new online reinforcement learning approach, enhances the self-correction ability of large language models, showing top performance with Gemini 1.0 Pro and 1.5 Flash models.
- [MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark](https://www.ml-quant.com/papers/arxiv/2406.01574/) (2024-10-09): MMLU-Pro, an improved dataset, expands the Massive Multitask Language Understanding benchmark by adding tougher questions and more choices, serving as a better benchmark to monitor progress in the field.
- [SaySelf: Teaching LLMs to Express Confidence with Self-Reflective Rationales](https://www.ml-quant.com/papers/arxiv/2405.20974/) (2024-10-09): The SaySelf training framework instructs large language models to provide more precise confidence estimates and self-reflective rationales, effectively reducing confidence calibration error while maintaining task performance.
- [FastCLIP: A Suite of Optimization Techniques to Accelerate CLIP Training with Limited Resources](https://www.ml-quant.com/papers/arxiv/2407.01445/) (2024-10-09): FastCLIP, a new CLIP training framework, optimizes the training of Contrastive Language-Image Pretraining models on large-scale data with limited resources, showing significant improvement in resource-limited settings.
- [AgentStudio: A Toolkit for Building General Virtual Agents](https://www.ml-quant.com/papers/arxiv/2403.17918/) (2024-10-09): AgentStudio, a combination of environments, tools, and benchmarks, aids in the development and evaluation of general virtual agents in real-world settings, offering a range of online tasks and datasets to assess fundamental agent abilities.
- [Accelerating Training with Neuron Interaction and Nowcasting Networks](https://www.ml-quant.com/papers/arxiv/2409.04434/) (2024-10-09): The article explores the enhancement of weight nowcaster networks (WNNs) through neuron interaction and nowcasting (NiNo) networks, resulting in a 50% speed increase in neural network training for vision and language tasks.
- [LML-DAP: LANGUAGE MODEL LEARNING A DATASET FOR DATA-AUGMENTED PREDICTION](https://www.ml-quant.com/papers/arxiv/2409.18957/) (2024-10-09): Language Model Prediction: The paper presents a new method called Data-Augmented Prediction (DAP) for using Large Language Models (LLMs) in classification tasks, achieving over 90% accuracy in some tests.
- [Scattering spectra models for physics](https://www.ml-quant.com/papers/arxiv/2306.17210/) (2024-10-09): The paper introduces scattering spectra models for stationary fields, offering precise and robust statistical descriptions for various physics fields, aiding in data exploration, classification, parameter inference, symmetry detection, and component separation.
- [SoK: Membership Inference Attacks on LLMs are Rushing Nowhere (and How to Fix It)](https://www.ml-quant.com/papers/arxiv/2406.17975/) (2024-10-09): The article examines Membership Inference Attacks (MIAs) against Large Language Models (LLMs), suggesting potential solutions and comprehensive benchmarks for sequence- and document-level MIAs against LLMs.
- [SwapAnything: Enabling Arbitrary Object Swapping in Personalized Image Editing](https://www.ml-quant.com/papers/arxiv/2404.05717/) (2024-10-09): Object Swapping: The paper presents SwapAnything, a new framework that allows for the swapping of any objects in an image with personalized concepts while maintaining the original context, showing significant improvement over previous methods in personalized swapping.
- [Unconditional stability of a recurrent neural circuit implementing divisive normalization](https://www.ml-quant.com/papers/arxiv/2409.18946/) (2024-10-03): The research introduces ORGaNICs, a recurrent cortical circuit model that outperforms other models in image classification tasks and matches LSTMs in sequential tasks, offering dynamic divisive normalization and unconditional local stability.
- [Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers](https://www.ml-quant.com/papers/arxiv/2409.20537/) (2024-10-03): The paper presents Heterogeneous Pre-trained Transformers (HPT), a method for training robotic models across various tasks, improving the performance of fine-tuned policies by over 20% on unseen tasks in both simulated and real-world environments.
- [What is the Role of Large Language Models in the Evolution of Astronomy Research?](https://www.ml-quant.com/papers/arxiv/2409.20252/) (2024-10-03): A study involving 13 astronomers discusses the potential and limitations of large language models like ChatGPT in research activities, emphasizing the importance of critical thinking and domain expertise to ensure these tools support, not replace, rigorous scientific investigation.
- [Maia-2: A Unified Model for Human-AI Alignment in Chess](https://www.ml-quant.com/papers/arxiv/2409.20553/) (2024-10-03): Researchers have proposed a unified model for aligning human and AI strategies in chess, which could lead to AI-based teaching tools.

All 836: https://www.ml-quant.com/api/v1/papers/ml.json
