Papers · 36 canon · 100 frontier

Papers

The research the Field Guide is built on: the all-time papers worth knowing, and last year's most-cited. Each one in plain words, with the ideas to read first.

01 · All time

The canon

36 papers that shaped the field, from the perceptron to Constitutional AI, one or more for every region of the map. In order of publication.

1958 Foundations

The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain

Frank Rosenblatt

It is the first learning neural network: the idea that weights, not rules, should hold the knowledge starts here.

1986 Training

Learning Representations by Back-Propagating Errors

David E. Rumelhart et al.

Backpropagation is how essentially every neural network, from LeNet to today’s LLMs, is trained.

1995 Foundations

Support-Vector Networks

Corinna Cortes et al.

SVMs were the default strong classifier for a decade and made margins and kernels part of every practitioner’s vocabulary.

1997 Neural Networks

Long Short-Term Memory

Sepp Hochreiter et al.

LSTMs powered speech recognition, translation and text generation for twenty years, until transformers took over.

1998 Neural Networks

Gradient-Based Learning Applied to Document Recognition

Yann LeCun et al.

It is the blueprint for the convolutional network, the architecture that later cracked computer vision.

2001 Foundations

Random Forests

Leo Breiman

It is still the first thing to try on tabular data, and the clearest demonstration that averaging diverse weak models beats one strong one.

2002 Evaluation

BLEU: a Method for Automatic Evaluation of Machine Translation

Kishore Papineni et al.

It showed that an automatic metric can drive a field’s progress, and it is still reported in translation papers.

2009 Evaluation

ImageNet: A Large-Scale Hierarchical Image Database

Jia Deng et al.

ImageNet and its annual challenge gave deep learning the benchmark it needed to prove itself.

2012 Training

Improving Neural Networks by Preventing Co-adaptation of Feature Detectors

Geoffrey E. Hinton et al.

Dropout became the standard regularizer of the deep-learning era and is still used inside large models.

2012 Neural Networks

ImageNet Classification with Deep Convolutional Neural Networks

Alex Krizhevsky et al.

This is the moment deep learning went mainstream: GPUs plus data plus depth beat decades of feature engineering.

2013 Language & LLMs

Efficient Estimation of Word Representations in Vector Space

Tomas Mikolov et al.

It made embeddings, meaning as geometry, the starting point of modern NLP.

2013 Agents & RL

Playing Atari with Deep Reinforcement Learning

Volodymyr Mnih et al.

It started deep reinforcement learning: one learner, raw perception, many tasks.

2014 Vision & Multimodal

Generative Adversarial Networks

Ian J. Goodfellow et al.

GANs launched the modern era of generated images and dominated image synthesis until diffusion arrived.

2014 Language & LLMs

Neural Machine Translation by Jointly Learning to Align and Translate

Dzmitry Bahdanau et al.

Attention was born here, three years before it became “all you need”.

2014 Training

Adam: A Method for Stochastic Optimization

Diederik P. Kingma et al.

Adam (and its descendant AdamW) is the default optimizer for training almost every modern network, LLMs included.

2015 Training

Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Sergey Ioffe et al.

It made very deep networks practical to train, and normalization layers became a standard part of the recipe.

2015 Shipping AI

Distilling the Knowledge in a Neural Network

Geoffrey Hinton et al.

Distillation is how big models become fast, cheap ones you can actually deploy.

2015 Neural Networks

Deep Residual Learning for Image Recognition

Kaiming He et al.

Residual connections are now in nearly every deep network, including every transformer.

2016 Agents & RL

Mastering the Game of Go with Deep Neural Networks and Tree Search

David Silver et al.

It proved that learned intuition plus search can master a problem long thought a decade away.

2017 Language & LLMs

Attention Is All You Need

Ashish Vaswani et al.

Every major LLM, and most modern vision and speech models, is a transformer.

2017 Agents & RL

Proximal Policy Optimization Algorithms

John Schulman et al.

PPO became the default RL algorithm, including for RLHF on the first ChatGPT-era models.

2018 Language & LLMs

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Jacob Devlin et al.

It established “pre-train once, fine-tune everywhere”, and still powers much of search and classification.

2020 Language & LLMs

Language Models are Few-Shot Learners

Tom B. Brown et al.

It showed that scale alone unlocks in-context learning, the capability today’s prompt-driven AI is built on.

2020 Language & LLMs

Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Patrick Lewis et al.

It named and defined the pattern behind most production LLM apps that answer from your own documents.

2020 Vision & Multimodal

Denoising Diffusion Probabilistic Models

Jonathan Ho et al.

DDPM is the recipe behind Stable Diffusion, DALL·E, Midjourney and today’s video models.

2020 Evaluation

Measuring Massive Multitask Language Understanding

Dan Hendrycks et al.

It became the headline benchmark in LLM release notes for years, and a case study in benchmarks saturating.

2020 Vision & Multimodal

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Alexey Dosovitskiy et al.

It brought vision and language onto one architecture, which made today’s multimodal models possible.

2021 Vision & Multimodal

Learning Transferable Visual Models From Natural Language Supervision

Alec Radford et al.

CLIP is the bridge between words and pictures inside text-to-image models and many vision-language systems.

2021 Training

LoRA: Low-Rank Adaptation of Large Language Models

Edward J. Hu et al.

It made adapting big models cheap enough for everyone, and is why fine-tunes are shared as small adapter files.

2021 Vision & Multimodal

High-Resolution Image Synthesis with Latent Diffusion Models

Robin Rombach et al.

It’s the architecture behind Stable Diffusion, which put high-quality text-to-image generation on ordinary GPUs.

2022 Language & LLMs

Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Jason Wei et al.

Thinking step by step went from a prompt trick to the basis of today’s reasoning models.

2022 Training

Training Compute-Optimal Large Language Models

Jordan Hoffmann et al.

It reset how labs size their models: data matters as much as parameter count.

2022 Shipping AI

Training Language Models to Follow Instructions with Human Feedback

Long Ouyang et al.

This recipe, RLHF, is what turned raw language models into assistants like ChatGPT.

2022 Shipping AI

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Tri Dao et al.

It is a big part of why long context windows became affordable, and it ships inside nearly every LLM stack.

2022 Agents & RL

ReAct: Synergizing Reasoning and Acting in Language Models

Shunyu Yao et al.

The think–act–observe loop in ReAct is the skeleton of nearly every LLM agent and coding harness today.

2022 Shipping AI

Constitutional AI: Harmlessness from AI Feedback

Yuntao Bai et al.

It showed that alignment can scale with AI feedback steered by explicit principles, and shaped how Claude is trained.

02 · Last year

The frontier: Aug 2025 – Aug 2026

The 100 most-cited AI papers of the year, ranked by citations as of Aug 9, 2026. Counts keep moving; ours are a snapshot.

Selection: 1kpapers.com by Together AI, most-cited as of Aug 9, 2026

  1. #1 1.8K citations Nov 2025 Vision & Multimodal

    Qwen3-VL Technical Report

    Shuai Bai et al.

    It is the most-cited paper of the year and a reference point for open multimodal models.

  2. #2 1.2K citations Aug 2025 Vision & Multimodal

    DINOv3

    Oriane Siméoni et al.

    A single frozen backbone that beats specialised models on dense vision tasks makes self-supervised vision a practical default.

  3. #3 1.2K citations Aug 2025 Vision & Multimodal

    InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

    Weiyun Wang et al.

    It narrows the gap between open and commercial multimodal models on reasoning and agent tasks.

  4. #4 875 citations Aug 2025 Vision & Multimodal

    Qwen-Image Technical Report

    Chenfei Wu et al.

    Legible, editable text in generated images was a long-standing weakness; this open model largely fixes it.

  5. #5 711 citations Nov 2025 Vision & Multimodal

    SAM 3: Segment Anything with Concepts

    Nicolas Carion et al.

    Segmentation moves from “click on the thing” to “name the thing”, doubling accuracy on this task.

  6. #6 671 citations Dec 2025 Language & LLMs

    DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

    DeepSeek-AI et al.

    It shows an open model competing with the best closed ones on reasoning while getting cheaper to run.

  7. #7 499 citations Nov 2025 Vision & Multimodal

    Depth Anything 3: Recovering the Visual Space from Any Views

    Haotong Lin et al.

    A simpler recipe beats specialised 3D pipelines, suggesting 3D perception can ride on general vision backbones.

  8. #8 409 citations Aug 2025 Language & LLMs

    GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

    GLM-4.5 Team et al.

    It is one of the strongest open models for agents and coding at a fraction of rivals’ active size.

  9. #9 405 citations Sep 2025 Vision & Multimodal

    Qwen3-Omni Technical Report

    Jin Xu et al.

    It shows a single open model can match specialist models across every modality at once.

  10. #10 313 citations Feb 2026 Agents & RL

    Kimi K2.5: Visual Agentic Intelligence

    Kimi Team et al.

    Parallel sub-agents trained into the model itself point to where agent orchestration is heading.

  11. #11 295 citations Feb 2026 Agents & RL

    GLM-5: from Vibe Coding to Agentic Engineering

    GLM-5-Team et al.

    It is a marker of open models moving from “vibe coding” to real agentic engineering.

  12. #12 282 citations Sep 2025 Language & LLMs

    Why Language Models Hallucinate

    Adam Tauman Kalai et al.

    It reframes hallucination as an incentive problem baked into how we measure models, not a mystery.

  13. #13 236 citations Oct 2025 Agents & RL

    Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models

    Qizheng Zhang et al.

    It gives a principled shape to “context engineering”, the craft behind most agent harnesses.

  14. #14 232 citations Dec 2025 Agents & RL

    Memory in the Age of AI Agents

    Yuyang Hu et al.

    A shared vocabulary for agent memory, one of the least settled parts of agent design.

  15. #15 226 citations Sep 2025 Vision & Multimodal

    Seedream 4.0: Toward Next-generation Multimodal Image Generation

    Team Seedream et al.

    It set the bar for commercial image generation and editing in one fast system.

  16. #16 221 citations Oct 2025 Vision & Multimodal

    Diffusion Transformers with Representation Autoencoders

    Boyang Zheng et al.

    It argues the generation and understanding sides of vision should share one latent space.

  17. #17 206 citations Nov 2025 Vision & Multimodal

    Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

    Z-Image Team et al.

    It shows state-of-the-art image generation does not require “scale at all costs”.

  18. #18 197 citations Nov 2025 Vision & Multimodal

    SAM 3D: 3Dfy Anything in Images

    SAM 3D Team et al.

    Breaking the 3D data bottleneck makes single-photo 3D practical for real images.

  19. #19 194 citations Sep 2025 Vision & Multimodal

    Video models are zero-shot learners and reasoners

    Thaddäus Wiedemer et al.

    It suggests video models may become general vision foundation models the way LLMs did for text.

  20. #20 189 citations Sep 2025 Vision & Multimodal

    LongLive: Real-time Interactive Long Video Generation

    Shuai Yang et al.

    Minute-long, steerable video at 20 frames per second on one GPU makes interactive video generation real.

  21. #21 182 citations Sep 2025 Agents & RL

    The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

    Guibin Zhang et al.

    The map to read first if you want to understand how agents are trained, not just prompted.

  22. #22 174 citations Feb 2026 Evaluation

    SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

    Xiangyi Li et al.

    It is the first careful evidence that skills work, and that smaller models with good skills can match bigger ones.

  23. #23 167 citations Aug 2025 Training

    R-Zero: Self-Evolving Reasoning LLM from Zero Data

    Chengsong Huang et al.

    Self-generated training data is one route around the human-data bottleneck.

  24. #24 167 citations Jan 2026 Vision & Multimodal

    Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking

    Mingxin Li et al.

    Good multimodal embeddings are what make RAG work over screenshots, PDFs and video.

  25. #25 163 citations Aug 2025 Language & LLMs

    Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference

    Yuxuan Song et al.

    It is evidence that diffusion language models can be dramatically faster without giving up quality.

  26. #26 163 citations Sep 2025 Agents & RL

    UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

    Haoming Wang et al.

    Computer-use agents are one of the fastest-moving frontiers; this report shows how one is trained at scale.

  27. #27 161 citations Sep 2025 Training

    A Survey of Reinforcement Learning for Large Reasoning Models

    Kaiyan Zhang et al.

    RL is now how reasoning is trained into models; this is the field’s reading list.

  28. #28 159 citations Oct 2025 Vision & Multimodal

    DeepSeek-OCR: Contexts Optical Compression

    Haoran Wei et al.

    It raises a new idea for long context: compress old history visually instead of keeping every text token.

  29. #29 156 citations Oct 2025 Vision & Multimodal

    Self-Forcing++: Towards Minute-Scale High-Quality Video Generation

    Justin Cui et al.

    It shows a way to long video without long training videos or long-video teachers.

  30. #30 154 citations Aug 2025 Agents & RL

    A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems

    Jinyuan Fang et al.

    Agents that keep learning on the job are a major research bet; this lays out the design space.

  31. #31 151 citations Aug 2025 Shipping AI

    Deep Think with Confidence

    Yichao Fu et al.

    Cheaper test-time reasoning with no extra training is directly useful to anyone serving reasoning models.

  32. #32 150 citations Feb 2026 Agents & RL

    SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning

    Peng Xia et al.

    It connects two hot ideas, agent skills and agentic RL, into one learning loop.

  33. #33 150 citations Aug 2025 Agents & RL

    Mobile-Agent-v3: Fundamental Agents for GUI Automation

    Jiabo Ye et al.

    A strong, open computer-use agent for researchers who can’t use closed ones.

  34. #34 145 citations Apr 2026 Training

    Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

    Yaxuan Li et al.

    Practical guidance on a post-training technique most labs now depend on.

  35. #35 139 citations Sep 2025 Training

    SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

    Zhenghai Xue et al.

    A one-idea fix that makes multi-turn tool-using RL trainable.

  36. #36 135 citations Jan 2026 Vision & Multimodal

    LTX-2: Efficient Joint Audio-Visual Foundation Model

    Yoav HaCohen et al.

    Generated video stops being silent, in an open model at a fraction of proprietary cost.

  37. #37 130 citations Sep 2025 Agents & RL

    SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning

    Haozhan Li et al.

    It brings the “RL after pre-training” recipe that transformed LLM reasoning to robot control.

  38. #38 129 citations Sep 2025 Agents & RL

    VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model

    Yihao Wang et al.

    It makes capable robot policies cheap enough for small labs.

  39. #39 124 citations Sep 2025 Vision & Multimodal

    MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe

    Tianyu Yu et al.

    Strong multimodal AI that runs on modest hardware.

  40. #40 123 citations Dec 2025 Language & LLMs

    LLaDA2.0: Scaling Up Diffusion Language Models to 100B

    Tiwei Bie et al.

    It shows diffusion LLMs can reach frontier scale by inheriting from existing models.

  41. #41 119 citations Aug 2025 Vision & Multimodal

    Thyme: Think Beyond Images

    Yi-Fan Zhang et al.

    An open take on “thinking with images”, a capability until now mainly seen in closed models.

  42. #42 116 citations Oct 2025 Neural Networks

    Kimi Linear: An Expressive, Efficient Attention Architecture

    Kimi Team et al.

    It is a credible claim that linear attention can replace full attention without a quality penalty.

  43. #43 114 citations Jan 2026 Training

    GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

    Shih-Yang Liu et al.

    Multi-objective RL is the norm in post-training; this fixes a subtle flaw in the default algorithm.

  44. #44 114 citations Sep 2025 Shipping AI

    Fast-dLLM v2: Efficient Block-Diffusion LLM

    Chengyue Wu et al.

    A cheap path to faster inference from models you already have.

  45. #45 113 citations Aug 2025 Agents & RL

    Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

    Yue Liao et al.

    It shows video generation doubling as the brain, simulator and test bench for robots.

  46. #46 112 citations Aug 2025 Training

    On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification

    Yongliang Wu et al.

    A tiny, theory-backed change to the most common training step in the stack.

  47. #47 111 citations Oct 2025 Agents & RL

    LightMem: Lightweight and Efficient Memory-Augmented Generation

    Jizhan Fang et al.

    Memory that is both better and far cheaper matters for any long-running assistant.

  48. #48 109 citations Oct 2025 Vision & Multimodal

    Emu3.5: Native Multimodal Models are World Learners

    Yufeng Cui et al.

    A strong case for one next-token model as a world learner across vision and language.

  49. #49 108 citations Oct 2025 Evaluation

    GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

    Tejal Patwardhan et al.

    It measures economic usefulness directly, instead of exam-style puzzles.

  50. #50 104 citations Oct 2025 Language & LLMs

    Scaling Latent Reasoning via Looped Language Models

    Rui-Jie Zhu et al.

    Looping depth is a new scaling knob for reasoning that small models can exploit.

  51. #51 101 citations Aug 2025 Agents & RL

    WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

    Xinyu Geng et al.

    Deep-research agents have been text-only; this extends them to the visual web.

  52. #52 101 citations Sep 2025 Vision & Multimodal

    Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search

    Xin Lai et al.

    It shows how to train long, exploratory visual reasoning in an open model.

  53. #53 97 citations Dec 2025 Vision & Multimodal

    WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling

    Wenqiang Sun et al.

    Interactive, consistent world simulation in real time edges generative video toward playable worlds.

  54. #54 95 citations Sep 2025 Agents & RL

    ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization

    Xixi Wu et al.

    A plug-in answer to the context limit every long-running agent eventually hits.

  55. #55 93 citations Oct 2025 Agents & RL

    Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model

    Fuhao Li et al.

    A cheap way to give robot policies a sense of 3D space.

  56. #56 92 citations Feb 2026 Agents & RL

    MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents

    Haozhen Zhang et al.

    Memory management that learns and evolves, rather than being hand-designed.

  57. #57 91 citations Aug 2025 Vision & Multimodal

    Self-Rewarding Vision-Language Model via Reasoning Decomposition

    Zongxia Li et al.

    It reduces visual hallucination without an external reward model.

  58. #58 90 citations Aug 2025 Vision & Multimodal

    Pref-GRPO: Pairwise Preference Reward-based GRPO for Stable Text-to-Image Reinforcement Learning

    Yibin Wang et al.

    More stable RL for image generators, plus a better yardstick for them.

  59. #59 90 citations Aug 2025 Evaluation

    MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP Servers

    Zhenting Wang et al.

    It measures tool use through MCP, the protocol agents actually use in practice.

  60. #60 88 citations Jan 2026 Vision & Multimodal

    Qwen3-TTS Technical Report

    Hangrui Hu et al.

    State-of-the-art, controllable open speech synthesis, released under Apache 2.0.

  61. #61 87 citations Aug 2025 Agents & RL

    Memento: Fine-tuning LLM Agents without Fine-tuning LLMs

    Huichi Zhou et al.

    Continual learning for agents without the cost of fine-tuning.

  62. #62 87 citations Jan 2026 Vision & Multimodal

    Advancing Open-source World Models

    Robbyant Team et al.

    It narrows the gap between open and closed interactive world models.

  63. #63 87 citations Feb 2026 Training

    Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

    Wenkai Yang et al.

    It turns a popular heuristic into a framework with a knob that measurably helps.

  64. #64 87 citations Apr 2026 Vision & Multimodal

    Qwen3.5-Omni Technical Report

    Qwen Team

    The frontier of open omni-modal models, rivalling Gemini on audio.

  65. #65 87 citations Oct 2025 Vision & Multimodal

    Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding

    Yi Xin et al.

    Evidence that diffusion can be the single engine for unified multimodal models.

  66. #66 86 citations Sep 2025 Vision & Multimodal

    MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing

    Junbo Niu et al.

    Accurate, cheap PDF parsing is the unglamorous foundation of document RAG.

  67. #67 82 citations Sep 2025 Vision & Multimodal

    SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer

    Junsong Chen et al.

    Quality video generation that fits on a single consumer GPU.

  68. #68 81 citations Aug 2025 Vision & Multimodal

    Ovis2.5 Technical Report

    Shiyin Lu et al.

    Small open multimodal models with strong reasoning, suited to on-device use.

  69. #69 81 citations Sep 2025 Agents & RL

    VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use

    Dongfu Jiang et al.

    Shared infrastructure for agentic RL, instead of one-off codebases per task.

  70. #70 78 citations Apr 2026 Vision & Multimodal

    Seedance 2.0: Advancing Video Generation for World Complexity

    Team Seedance et al.

    A commercial frontier for controllable audio-video generation.

  71. #71 78 citations Mar 2026 Agents & RL

    Trace2Skill: Distill Trajectory-Local Lessons into Transferable Agent Skills

    Jingwei Ni et al.

    Automatic skill writing from experience, with skills that travel between models.

  72. #72 76 citations Oct 2025 Shipping AI

    Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples

    Alexandra Souly et al.

    Bigger training sets do not dilute poisoning, so attacks may be easier than assumed.

  73. #73 75 citations Nov 2025 Agents & RL

    DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

    Rulin Shao et al.

    It shows how to use RL where there is no verifiable answer.

  74. #74 75 citations Sep 2025 Training

    Revolutionizing Reinforcement Learning Framework for Diffusion Large Language Models

    Yinjie Wang et al.

    It brings the RL post-training playbook to diffusion LLMs.

  75. #75 74 citations Sep 2025 Agents & RL

    WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents

    Junteng Liu et al.

    Hard synthetic data, not model size, turns out to be the lever for web agents.

  76. #76 73 citations Oct 2025 Vision & Multimodal

    StreamingVLM: Real-Time Understanding for Infinite Video Streams

    Ruyi Xu et al.

    Real-time assistants that watch a stream need exactly this kind of bounded-memory attention.

  77. #77 71 citations Jan 2026 Training

    Learning to Discover at Test Time

    Mert Yuksekgonul et al.

    Test-time training as a tool for scientific discovery, reproducible with open models.

  78. #78 69 citations Oct 2025 Vision & Multimodal

    PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model

    Cheng Cui et al.

    Tiny, multilingual and accurate document parsing ready for production.

  79. #79 69 citations Oct 2025 Vision & Multimodal

    The Principles of Diffusion Models

    Chieh-Hsin Lai et al.

    One coherent mental model for the math behind image, video and diffusion language models.

  80. #80 68 citations Oct 2025 Vision & Multimodal

    UniVideo: Unified Understanding, Generation, and Editing for Videos

    Cong Wei et al.

    One model for understanding, generating and editing video, instead of a tool per task.

  81. #81 68 citations Mar 2026 Neural Networks

    Mamba-3: Improved Sequence Modeling using State Space Principles

    Aakash Lahoti et al.

    Inference cost now dominates, which makes fast, constant-memory architectures matter again.

  82. #82 66 citations Mar 2026 Agents & RL

    OpenClaw-RL: Train Any Agent Simply by Talking

    Yinjie Wang et al.

    Agents that learn continuously from ordinary use, without a separate training phase.

  83. #83 66 citations Aug 2025 Agents & RL

    Chain-of-Agents: End-to-End Agent Foundation Models via Multi-Agent Distillation and Agentic RL

    Weizhen Li et al.

    It asks whether multi-agent orchestration can be folded into a single model.

  84. #84 66 citations Sep 2025 Agents & RL

    WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents

    Zile Qiao et al.

    A clear pattern for keeping long research agents from drowning in their own context.

  85. #85 65 citations Dec 2025 Neural Networks

    mHC: Manifold-Constrained Hyper-Connections

    Zhenda Xie et al.

    The residual connection hadn’t changed in a decade; this is a stable way to go beyond it.

  86. #86 65 citations Oct 2025 Agents & RL

    AgentFold: Long-Horizon Web Agents with Proactive Context Management

    Rui Ye et al.

    Learned context management as an alternative to ever-bigger context windows.

  87. #87 65 citations Sep 2025 Vision & Multimodal

    3D and 4D World Modeling: A Survey

    Lingdong Kong et al.

    “World model” means many things; this pins down the 3D sense of it.

  88. #88 64 citations Dec 2025 Language & LLMs

    Recursive Language Models

    Alex L. Zhang et al.

    A different answer to long context: let the model navigate the input instead of reading it all.

  89. #89 64 citations Aug 2025 Agents & RL

    Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory

    Lin Long et al.

    A step toward assistants that remember what they have seen and heard over time.

  90. #90 64 citations Nov 2025 Vision & Multimodal

    Visual Spatial Tuning

    Rui Yang et al.

    Spatial reasoning is a known weak spot of VLMs; this fixes much of it with data, not extra encoders.

  91. #91 63 citations Oct 2025 Agents & RL

    DeepAgent: A General Reasoning Agent with Scalable Toolsets

    Xiaoxi Li et al.

    A general agent design that scales to large, open-ended toolsets.

  92. #92 63 citations Feb 2026 Shipping AI

    DFlash: Block Diffusion for Flash Speculative Decoding

    Jian Chen et al.

    Faster LLM serving with identical outputs, combining diffusion drafting with standard models.

  93. #93 63 citations Jan 2026 Neural Networks

    Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

    Xin Cheng et al.

    It proposes memory lookup as a new scaling axis for LLMs, separate from compute.

  94. #94 62 citations Nov 2025 Agents & RL

    MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

    MiroMind Team et al.

    It frames “more interaction” as scaling, alongside bigger models and longer contexts.

  95. #95 62 citations Sep 2025 Agents & RL

    AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

    Zhiheng Xi et al.

    A reusable gym for multi-turn agent RL, with agents that rival commercial models on 27 tasks.

  96. #96 61 citations Mar 2026 Training

    Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?

    Jeonghye Kim et al.

    Shorter reasoning isn’t free: saying “I’m not sure” is part of how models reason well.

  97. #97 60 citations Nov 2025 Training

    DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning

    Zhihong Shao et al.

    Self-verification is a path to reasoning where no answer key exists.

  98. #98 60 citations Sep 2025 Agents & RL

    WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning

    Kuan Li et al.

    A detailed open recipe for the hardest information-seeking tasks.

  99. #99 60 citations Oct 2025 Vision & Multimodal

    Detect Anything via Next Point Prediction

    Qing Jiang et al.

    Language models are catching up to specialised detectors while adding pointing, referring and OCR.

  100. #100 59 citations Aug 2025 Training

    TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

    Yizhi Li et al.

    RL post-training is expensive; sharing work across rollouts makes it cheaper.

Keep up

The papers worth your time, explained.

New paper notes, Field Guide deep dives and lessons, one email at a time. Free, no spam.