Papers · 36 canon · 100 frontier
Papers
The research the Field Guide is built on: the all-time papers worth knowing, and last year's most-cited. Each one in plain words, with the ideas to read first.
01 · All time
The canon
36 papers that shaped the field, from the perceptron to Constitutional AI, one or more for every region of the map. In order of publication.
1958 Foundations
The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain
Frank Rosenblatt
It is the first learning neural network: the idea that weights, not rules, should hold the knowledge starts here.
1986 Training
Learning Representations by Back-Propagating Errors
David E. Rumelhart et al.
Backpropagation is how essentially every neural network, from LeNet to today’s LLMs, is trained.
1995 Foundations
Support-Vector Networks
Corinna Cortes et al.
SVMs were the default strong classifier for a decade and made margins and kernels part of every practitioner’s vocabulary.
1997 Neural Networks
Long Short-Term Memory
Sepp Hochreiter et al.
LSTMs powered speech recognition, translation and text generation for twenty years, until transformers took over.
1998 Neural Networks
Gradient-Based Learning Applied to Document Recognition
Yann LeCun et al.
It is the blueprint for the convolutional network, the architecture that later cracked computer vision.
2001 Foundations
Random Forests
Leo Breiman
It is still the first thing to try on tabular data, and the clearest demonstration that averaging diverse weak models beats one strong one.
2002 Evaluation
BLEU: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni et al.
It showed that an automatic metric can drive a field’s progress, and it is still reported in translation papers.
2009 Evaluation
ImageNet: A Large-Scale Hierarchical Image Database
Jia Deng et al.
ImageNet and its annual challenge gave deep learning the benchmark it needed to prove itself.
2012 Training
Improving Neural Networks by Preventing Co-adaptation of Feature Detectors
Geoffrey E. Hinton et al.
Dropout became the standard regularizer of the deep-learning era and is still used inside large models.
2012 Neural Networks
ImageNet Classification with Deep Convolutional Neural Networks
Alex Krizhevsky et al.
This is the moment deep learning went mainstream: GPUs plus data plus depth beat decades of feature engineering.
2013 Language & LLMs
Efficient Estimation of Word Representations in Vector Space
Tomas Mikolov et al.
It made embeddings, meaning as geometry, the starting point of modern NLP.
2013 Agents & RL
Playing Atari with Deep Reinforcement Learning
Volodymyr Mnih et al.
It started deep reinforcement learning: one learner, raw perception, many tasks.
2014 Vision & Multimodal
Generative Adversarial Networks
Ian J. Goodfellow et al.
GANs launched the modern era of generated images and dominated image synthesis until diffusion arrived.
2014 Language & LLMs
Neural Machine Translation by Jointly Learning to Align and Translate
Dzmitry Bahdanau et al.
Attention was born here, three years before it became “all you need”.
2014 Training
Adam: A Method for Stochastic Optimization
Diederik P. Kingma et al.
Adam (and its descendant AdamW) is the default optimizer for training almost every modern network, LLMs included.
2015 Training
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Sergey Ioffe et al.
It made very deep networks practical to train, and normalization layers became a standard part of the recipe.
2015 Shipping AI
Distilling the Knowledge in a Neural Network
Geoffrey Hinton et al.
Distillation is how big models become fast, cheap ones you can actually deploy.
2015 Neural Networks
Deep Residual Learning for Image Recognition
Kaiming He et al.
Residual connections are now in nearly every deep network, including every transformer.
2016 Agents & RL
Mastering the Game of Go with Deep Neural Networks and Tree Search
David Silver et al.
It proved that learned intuition plus search can master a problem long thought a decade away.
2017 Language & LLMs
Attention Is All You Need
Ashish Vaswani et al.
Every major LLM, and most modern vision and speech models, is a transformer.
2017 Agents & RL
Proximal Policy Optimization Algorithms
John Schulman et al.
PPO became the default RL algorithm, including for RLHF on the first ChatGPT-era models.
2018 Language & LLMs
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin et al.
It established “pre-train once, fine-tune everywhere”, and still powers much of search and classification.
2020 Language & LLMs
Language Models are Few-Shot Learners
Tom B. Brown et al.
It showed that scale alone unlocks in-context learning, the capability today’s prompt-driven AI is built on.
2020 Language & LLMs
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis et al.
It named and defined the pattern behind most production LLM apps that answer from your own documents.
2020 Vision & Multimodal
Denoising Diffusion Probabilistic Models
Jonathan Ho et al.
DDPM is the recipe behind Stable Diffusion, DALL·E, Midjourney and today’s video models.
2020 Evaluation
Measuring Massive Multitask Language Understanding
Dan Hendrycks et al.
It became the headline benchmark in LLM release notes for years, and a case study in benchmarks saturating.
2020 Vision & Multimodal
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy et al.
It brought vision and language onto one architecture, which made today’s multimodal models possible.
2021 Vision & Multimodal
Learning Transferable Visual Models From Natural Language Supervision
Alec Radford et al.
CLIP is the bridge between words and pictures inside text-to-image models and many vision-language systems.
2021 Training
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu et al.
It made adapting big models cheap enough for everyone, and is why fine-tunes are shared as small adapter files.
2021 Vision & Multimodal
High-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach et al.
It’s the architecture behind Stable Diffusion, which put high-quality text-to-image generation on ordinary GPUs.
2022 Language & LLMs
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Jason Wei et al.
Thinking step by step went from a prompt trick to the basis of today’s reasoning models.
2022 Training
Training Compute-Optimal Large Language Models
Jordan Hoffmann et al.
It reset how labs size their models: data matters as much as parameter count.
2022 Shipping AI
Training Language Models to Follow Instructions with Human Feedback
Long Ouyang et al.
This recipe, RLHF, is what turned raw language models into assistants like ChatGPT.
2022 Shipping AI
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Tri Dao et al.
It is a big part of why long context windows became affordable, and it ships inside nearly every LLM stack.
2022 Agents & RL
ReAct: Synergizing Reasoning and Acting in Language Models
Shunyu Yao et al.
The think–act–observe loop in ReAct is the skeleton of nearly every LLM agent and coding harness today.
2022 Shipping AI
Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai et al.
It showed that alignment can scale with AI feedback steered by explicit principles, and shaped how Claude is trained.
02 · Last year
The frontier: Aug 2025 – Aug 2026
The 100 most-cited AI papers of the year, ranked by citations as of Aug 9, 2026. Counts keep moving; ours are a snapshot.
Selection: 1kpapers.com by Together AI, most-cited as of Aug 9, 2026
-
#1 1.8K citations Nov 2025 Vision & Multimodal
Qwen3-VL Technical Report
Shuai Bai et al.
It is the most-cited paper of the year and a reference point for open multimodal models.
-
#2 1.2K citations Aug 2025 Vision & Multimodal
DINOv3
Oriane Siméoni et al.
A single frozen backbone that beats specialised models on dense vision tasks makes self-supervised vision a practical default.
-
#3 1.2K citations Aug 2025 Vision & Multimodal
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Weiyun Wang et al.
It narrows the gap between open and commercial multimodal models on reasoning and agent tasks.
-
#4 875 citations Aug 2025 Vision & Multimodal
Qwen-Image Technical Report
Chenfei Wu et al.
Legible, editable text in generated images was a long-standing weakness; this open model largely fixes it.
-
#5 711 citations Nov 2025 Vision & Multimodal
SAM 3: Segment Anything with Concepts
Nicolas Carion et al.
Segmentation moves from “click on the thing” to “name the thing”, doubling accuracy on this task.
-
#6 671 citations Dec 2025 Language & LLMs
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
DeepSeek-AI et al.
It shows an open model competing with the best closed ones on reasoning while getting cheaper to run.
-
#7 499 citations Nov 2025 Vision & Multimodal
Depth Anything 3: Recovering the Visual Space from Any Views
Haotong Lin et al.
A simpler recipe beats specialised 3D pipelines, suggesting 3D perception can ride on general vision backbones.
-
#8 409 citations Aug 2025 Language & LLMs
GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
GLM-4.5 Team et al.
It is one of the strongest open models for agents and coding at a fraction of rivals’ active size.
-
#9 405 citations Sep 2025 Vision & Multimodal
Qwen3-Omni Technical Report
Jin Xu et al.
It shows a single open model can match specialist models across every modality at once.
-
#10 313 citations Feb 2026 Agents & RL
Kimi K2.5: Visual Agentic Intelligence
Kimi Team et al.
Parallel sub-agents trained into the model itself point to where agent orchestration is heading.
-
#11 295 citations Feb 2026 Agents & RL
GLM-5: from Vibe Coding to Agentic Engineering
GLM-5-Team et al.
It is a marker of open models moving from “vibe coding” to real agentic engineering.
-
#12 282 citations Sep 2025 Language & LLMs
Why Language Models Hallucinate
Adam Tauman Kalai et al.
It reframes hallucination as an incentive problem baked into how we measure models, not a mystery.
-
#13 236 citations Oct 2025 Agents & RL
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
Qizheng Zhang et al.
It gives a principled shape to “context engineering”, the craft behind most agent harnesses.
-
#14 232 citations Dec 2025 Agents & RL
Memory in the Age of AI Agents
Yuyang Hu et al.
A shared vocabulary for agent memory, one of the least settled parts of agent design.
-
#15 226 citations Sep 2025 Vision & Multimodal
Seedream 4.0: Toward Next-generation Multimodal Image Generation
Team Seedream et al.
It set the bar for commercial image generation and editing in one fast system.
-
#16 221 citations Oct 2025 Vision & Multimodal
Diffusion Transformers with Representation Autoencoders
Boyang Zheng et al.
It argues the generation and understanding sides of vision should share one latent space.
-
#17 206 citations Nov 2025 Vision & Multimodal
Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
Z-Image Team et al.
It shows state-of-the-art image generation does not require “scale at all costs”.
-
#18 197 citations Nov 2025 Vision & Multimodal
SAM 3D: 3Dfy Anything in Images
SAM 3D Team et al.
Breaking the 3D data bottleneck makes single-photo 3D practical for real images.
-
#19 194 citations Sep 2025 Vision & Multimodal
Video models are zero-shot learners and reasoners
Thaddäus Wiedemer et al.
It suggests video models may become general vision foundation models the way LLMs did for text.
-
#20 189 citations Sep 2025 Vision & Multimodal
LongLive: Real-time Interactive Long Video Generation
Shuai Yang et al.
Minute-long, steerable video at 20 frames per second on one GPU makes interactive video generation real.
-
#21 182 citations Sep 2025 Agents & RL
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Guibin Zhang et al.
The map to read first if you want to understand how agents are trained, not just prompted.
-
#22 174 citations Feb 2026 Evaluation
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Xiangyi Li et al.
It is the first careful evidence that skills work, and that smaller models with good skills can match bigger ones.
-
#23 167 citations Aug 2025 Training
R-Zero: Self-Evolving Reasoning LLM from Zero Data
Chengsong Huang et al.
Self-generated training data is one route around the human-data bottleneck.
-
#24 167 citations Jan 2026 Vision & Multimodal
Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
Mingxin Li et al.
Good multimodal embeddings are what make RAG work over screenshots, PDFs and video.
-
#25 163 citations Aug 2025 Language & LLMs
Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference
Yuxuan Song et al.
It is evidence that diffusion language models can be dramatically faster without giving up quality.
-
#26 163 citations Sep 2025 Agents & RL
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
Haoming Wang et al.
Computer-use agents are one of the fastest-moving frontiers; this report shows how one is trained at scale.
-
#27 161 citations Sep 2025 Training
A Survey of Reinforcement Learning for Large Reasoning Models
Kaiyan Zhang et al.
RL is now how reasoning is trained into models; this is the field’s reading list.
-
#28 159 citations Oct 2025 Vision & Multimodal
DeepSeek-OCR: Contexts Optical Compression
Haoran Wei et al.
It raises a new idea for long context: compress old history visually instead of keeping every text token.
-
#29 156 citations Oct 2025 Vision & Multimodal
Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
Justin Cui et al.
It shows a way to long video without long training videos or long-video teachers.
-
#30 154 citations Aug 2025 Agents & RL
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
Jinyuan Fang et al.
Agents that keep learning on the job are a major research bet; this lays out the design space.
-
#31 151 citations Aug 2025 Shipping AI
Deep Think with Confidence
Yichao Fu et al.
Cheaper test-time reasoning with no extra training is directly useful to anyone serving reasoning models.
-
#32 150 citations Feb 2026 Agents & RL
SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning
Peng Xia et al.
It connects two hot ideas, agent skills and agentic RL, into one learning loop.
-
#33 150 citations Aug 2025 Agents & RL
Mobile-Agent-v3: Fundamental Agents for GUI Automation
Jiabo Ye et al.
A strong, open computer-use agent for researchers who can’t use closed ones.
-
#34 145 citations Apr 2026 Training
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
Yaxuan Li et al.
Practical guidance on a post-training technique most labs now depend on.
-
#35 139 citations Sep 2025 Training
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning
Zhenghai Xue et al.
A one-idea fix that makes multi-turn tool-using RL trainable.
-
#36 135 citations Jan 2026 Vision & Multimodal
LTX-2: Efficient Joint Audio-Visual Foundation Model
Yoav HaCohen et al.
Generated video stops being silent, in an open model at a fraction of proprietary cost.
-
#37 130 citations Sep 2025 Agents & RL
SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
Haozhan Li et al.
It brings the “RL after pre-training” recipe that transformed LLM reasoning to robot control.
-
#38 129 citations Sep 2025 Agents & RL
VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
Yihao Wang et al.
It makes capable robot policies cheap enough for small labs.
-
#39 124 citations Sep 2025 Vision & Multimodal
MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe
Tianyu Yu et al.
Strong multimodal AI that runs on modest hardware.
-
#40 123 citations Dec 2025 Language & LLMs
LLaDA2.0: Scaling Up Diffusion Language Models to 100B
Tiwei Bie et al.
It shows diffusion LLMs can reach frontier scale by inheriting from existing models.
-
#41 119 citations Aug 2025 Vision & Multimodal
Thyme: Think Beyond Images
Yi-Fan Zhang et al.
An open take on “thinking with images”, a capability until now mainly seen in closed models.
-
#42 116 citations Oct 2025 Neural Networks
Kimi Linear: An Expressive, Efficient Attention Architecture
Kimi Team et al.
It is a credible claim that linear attention can replace full attention without a quality penalty.
-
#43 114 citations Jan 2026 Training
GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Shih-Yang Liu et al.
Multi-objective RL is the norm in post-training; this fixes a subtle flaw in the default algorithm.
-
#44 114 citations Sep 2025 Shipping AI
Fast-dLLM v2: Efficient Block-Diffusion LLM
Chengyue Wu et al.
A cheap path to faster inference from models you already have.
-
#45 113 citations Aug 2025 Agents & RL
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
Yue Liao et al.
It shows video generation doubling as the brain, simulator and test bench for robots.
-
#46 112 citations Aug 2025 Training
On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
Yongliang Wu et al.
A tiny, theory-backed change to the most common training step in the stack.
-
#47 111 citations Oct 2025 Agents & RL
LightMem: Lightweight and Efficient Memory-Augmented Generation
Jizhan Fang et al.
Memory that is both better and far cheaper matters for any long-running assistant.
-
#48 109 citations Oct 2025 Vision & Multimodal
Emu3.5: Native Multimodal Models are World Learners
Yufeng Cui et al.
A strong case for one next-token model as a world learner across vision and language.
-
#49 108 citations Oct 2025 Evaluation
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
Tejal Patwardhan et al.
It measures economic usefulness directly, instead of exam-style puzzles.
-
#50 104 citations Oct 2025 Language & LLMs
Scaling Latent Reasoning via Looped Language Models
Rui-Jie Zhu et al.
Looping depth is a new scaling knob for reasoning that small models can exploit.
-
#51 101 citations Aug 2025 Agents & RL
WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent
Xinyu Geng et al.
Deep-research agents have been text-only; this extends them to the visual web.
-
#52 101 citations Sep 2025 Vision & Multimodal
Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
Xin Lai et al.
It shows how to train long, exploratory visual reasoning in an open model.
-
#53 97 citations Dec 2025 Vision & Multimodal
WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
Wenqiang Sun et al.
Interactive, consistent world simulation in real time edges generative video toward playable worlds.
-
#54 95 citations Sep 2025 Agents & RL
ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization
Xixi Wu et al.
A plug-in answer to the context limit every long-running agent eventually hits.
-
#55 93 citations Oct 2025 Agents & RL
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
Fuhao Li et al.
A cheap way to give robot policies a sense of 3D space.
-
#56 92 citations Feb 2026 Agents & RL
MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
Haozhen Zhang et al.
Memory management that learns and evolves, rather than being hand-designed.
-
#57 91 citations Aug 2025 Vision & Multimodal
Self-Rewarding Vision-Language Model via Reasoning Decomposition
Zongxia Li et al.
It reduces visual hallucination without an external reward model.
-
#58 90 citations Aug 2025 Vision & Multimodal
Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning
Yibin Wang et al.
More stable RL for image generators, plus a better yardstick for them.
-
#59 90 citations Aug 2025 Evaluation
MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers
Zhenting Wang et al.
It measures tool use through MCP, the protocol agents actually use in practice.
-
#60 88 citations Jan 2026 Vision & Multimodal
Qwen3-TTS Technical Report
Hangrui Hu et al.
State-of-the-art, controllable open speech synthesis, released under Apache 2.0.
-
#61 87 citations Aug 2025 Agents & RL
Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
Huichi Zhou et al.
Continual learning for agents without the cost of fine-tuning.
-
#62 87 citations Jan 2026 Vision & Multimodal
Advancing Open-source World Models
Robbyant Team et al.
It narrows the gap between open and closed interactive world models.
-
#63 87 citations Feb 2026 Training
Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
Wenkai Yang et al.
It turns a popular heuristic into a framework with a knob that measurably helps.
-
#64 87 citations Apr 2026 Vision & Multimodal
Qwen3.5-Omni Technical Report
Qwen Team
The frontier of open omni-modal models, rivalling Gemini on audio.
-
#65 87 citations Oct 2025 Vision & Multimodal
Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
Yi Xin et al.
Evidence that diffusion can be the single engine for unified multimodal models.
-
#66 86 citations Sep 2025 Vision & Multimodal
MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
Junbo Niu et al.
Accurate, cheap PDF parsing is the unglamorous foundation of document RAG.
-
#67 82 citations Sep 2025 Vision & Multimodal
SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
Junsong Chen et al.
Quality video generation that fits on a single consumer GPU.
-
#68 81 citations Aug 2025 Vision & Multimodal
Ovis2.5 Technical Report
Shiyin Lu et al.
Small open multimodal models with strong reasoning, suited to on-device use.
-
#69 81 citations Sep 2025 Agents & RL
VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
Dongfu Jiang et al.
Shared infrastructure for agentic RL, instead of one-off codebases per task.
-
#70 78 citations Apr 2026 Vision & Multimodal
Seedance 2.0: Advancing Video Generation for World Complexity
Team Seedance et al.
A commercial frontier for controllable audio-video generation.
-
#71 78 citations Mar 2026 Agents & RL
Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills
Jingwei Ni et al.
Automatic skill writing from experience, with skills that travel between models.
-
#72 76 citations Oct 2025 Shipping AI
Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
Alexandra Souly et al.
Bigger training sets do not dilute poisoning, so attacks may be easier than assumed.
-
#73 75 citations Nov 2025 Agents & RL
DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Rulin Shao et al.
It shows how to use RL where there is no verifiable answer.
-
#74 75 citations Sep 2025 Training
Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models
Yinjie Wang et al.
It brings the RL post-training playbook to diffusion LLMs.
-
#75 74 citations Sep 2025 Agents & RL
WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
Junteng Liu et al.
Hard synthetic data, not model size, turns out to be the lever for web agents.
-
#76 73 citations Oct 2025 Vision & Multimodal
StreamingVLM: Real-Time Understanding for Infinite Video Streams
Ruyi Xu et al.
Real-time assistants that watch a stream need exactly this kind of bounded-memory attention.
-
#77 71 citations Jan 2026 Training
Learning to Discover at Test Time
Mert Yuksekgonul et al.
Test-time training as a tool for scientific discovery, reproducible with open models.
-
#78 69 citations Oct 2025 Vision & Multimodal
PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
Cheng Cui et al.
Tiny, multilingual and accurate document parsing ready for production.
-
#79 69 citations Oct 2025 Vision & Multimodal
The Principles of Diffusion Models
Chieh-Hsin Lai et al.
One coherent mental model for the math behind image, video and diffusion language models.
-
#80 68 citations Oct 2025 Vision & Multimodal
UniVideo: Unified Understanding, Generation, and Editing for Videos
Cong Wei et al.
One model for understanding, generating and editing video, instead of a tool per task.
-
#81 68 citations Mar 2026 Neural Networks
Mamba-3: Improved Sequence Modeling using State Space Principles
Aakash Lahoti et al.
Inference cost now dominates, which makes fast, constant-memory architectures matter again.
-
#82 66 citations Mar 2026 Agents & RL
OpenClaw-RL: Train Any Agent Simply by Talking
Yinjie Wang et al.
Agents that learn continuously from ordinary use, without a separate training phase.
-
#83 66 citations Aug 2025 Agents & RL
Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL
Weizhen Li et al.
It asks whether multi-agent orchestration can be folded into a single model.
-
#84 66 citations Sep 2025 Agents & RL
WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
Zile Qiao et al.
A clear pattern for keeping long research agents from drowning in their own context.
-
#85 65 citations Dec 2025 Neural Networks
mHC: Manifold-Constrained Hyper-Connections
Zhenda Xie et al.
The residual connection hadn’t changed in a decade; this is a stable way to go beyond it.
-
#86 65 citations Oct 2025 Agents & RL
AgentFold: Long-Horizon Web Agents with Proactive Context Management
Rui Ye et al.
Learned context management as an alternative to ever-bigger context windows.
-
#87 65 citations Sep 2025 Vision & Multimodal
3D and 4D World Modeling: A Survey
Lingdong Kong et al.
“World model” means many things; this pins down the 3D sense of it.
-
#88 64 citations Dec 2025 Language & LLMs
Recursive Language Models
Alex L. Zhang et al.
A different answer to long context: let the model navigate the input instead of reading it all.
-
#89 64 citations Aug 2025 Agents & RL
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
Lin Long et al.
A step toward assistants that remember what they have seen and heard over time.
-
#90 64 citations Nov 2025 Vision & Multimodal
Visual Spatial Tuning
Rui Yang et al.
Spatial reasoning is a known weak spot of VLMs; this fixes much of it with data, not extra encoders.
-
#91 63 citations Oct 2025 Agents & RL
DeepAgent: A General Reasoning Agent with Scalable Toolsets
Xiaoxi Li et al.
A general agent design that scales to large, open-ended toolsets.
-
#92 63 citations Feb 2026 Shipping AI
DFlash: Block Diffusion for Flash Speculative Decoding
Jian Chen et al.
Faster LLM serving with identical outputs, combining diffusion drafting with standard models.
-
#93 63 citations Jan 2026 Neural Networks
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
Xin Cheng et al.
It proposes memory lookup as a new scaling axis for LLMs, separate from compute.
-
#94 62 citations Nov 2025 Agents & RL
MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
MiroMind Team et al.
It frames “more interaction” as scaling, alongside bigger models and longer contexts.
-
#95 62 citations Sep 2025 Agents & RL
AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
Zhiheng Xi et al.
A reusable gym for multi-turn agent RL, with agents that rival commercial models on 27 tasks.
-
#96 61 citations Mar 2026 Training
Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
Jeonghye Kim et al.
Shorter reasoning isn’t free: saying “I’m not sure” is part of how models reason well.
-
#97 60 citations Nov 2025 Training
DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning
Zhihong Shao et al.
Self-verification is a path to reasoning where no answer key exists.
-
#98 60 citations Sep 2025 Agents & RL
WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning
Kuan Li et al.
A detailed open recipe for the hardest information-seeking tasks.
-
#99 60 citations Oct 2025 Vision & Multimodal
Detect Anything via Next Point Prediction
Qing Jiang et al.
Language models are catching up to specialised detectors while adding pointing, referring and OCR.
-
#100 59 citations Aug 2025 Training
TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Yizhi Li et al.
RL post-training is expensive; sharing work across rollouts makes it cheaper.
Keep up
The papers worth your time, explained.
New paper notes, Field Guide deep dives and lessons, one email at a time. Free, no spam.
Check your inbox to confirm.