| A | B | C | D | E | F | G | H | I | J | K | L | M | N | O | P | Q | R | S | T | U | V | W | X | Y | Z | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
1 | Date | Topic | Papers | |||||||||||||||||||||||
2 | 2/2/2026 | Pretraining | Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization | |||||||||||||||||||||||
3 | Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning | |||||||||||||||||||||||||
4 | 2/4/2026 | Embeddings | MMTEB: Massive Multilingual Text Embedding Benchmark | |||||||||||||||||||||||
5 | Improving Text Embeddings with Large Language Models | |||||||||||||||||||||||||
6 | 2/9/2026 | Embeddings | DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model (only Section 2.1) | |||||||||||||||||||||||
7 | NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models | |||||||||||||||||||||||||
8 | 2/11/2026 | Mixture-of-Experts | Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer | |||||||||||||||||||||||
9 | Your Mixture-of-Experts LLM Is Secretly an Embedding Model for Free | |||||||||||||||||||||||||
10 | 2/16/2026 | Reasoning | BIRD: A Trustworthy Bayesian Inference Framework for Large Language Models | |||||||||||||||||||||||
11 | 2/18/2026 | Reasoning / Alignment | Direct Preference Optimization: Your Language Model is Secretly a Reward Model | |||||||||||||||||||||||
12 | Training a Generally Curious Agent | |||||||||||||||||||||||||
13 | 2/23/2026 | Agent Harness | Agent Workflow Memory | |||||||||||||||||||||||
14 | The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents | |||||||||||||||||||||||||
15 | 2/25/2026 | Long-Context | Efficient Streaming Language Models with Attention Sinks | |||||||||||||||||||||||
16 | 3/2/2026 | Latent Attention | TransMLA: Multi-Head Latent Attention Is All You Need | |||||||||||||||||||||||
17 | DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models (only Section 2) | |||||||||||||||||||||||||
18 | 3/4/2026 | Reinforcement Learning | DAPO: An Open-Source LLM Reinforcement Learning System at Scale | |||||||||||||||||||||||
19 | Understanding R1-Zero-Like Training | |||||||||||||||||||||||||
20 | 3/9/2026 | Reasoning | Optimas: Optimizing Compound AI Systems with Globally Aligned Local Rewards | |||||||||||||||||||||||
21 | 3/11/2026 | Reasoning | CWM: An Open-Weights LLM for Research on Code Generation with World Models | |||||||||||||||||||||||
22 | 3/16/2026 | Self-Play RL | Absolute Zero: Reinforced Self-play Reasoning with Zero Data | |||||||||||||||||||||||
23 | 3/18/2026 | Self-Play RL | SPICE: Self-Play In Corpus Environments Improves Reasoning | |||||||||||||||||||||||
24 | Toward Training Superintelligent Software Agents through Self-Play SWE-RL | |||||||||||||||||||||||||
25 | 3/30/2026 | Test-Time Scaling | Scaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning | |||||||||||||||||||||||
26 | Learning to Discover at Test Time | |||||||||||||||||||||||||
27 | 4/1/2026 | Mode Collapse | Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond) | |||||||||||||||||||||||
28 | 4/6/2026 | Linear Transformer | Linear transformers are secretly fast weight programmers | |||||||||||||||||||||||
29 | Parallelizing Linear Transformers with the Delta Rule over Sequence Length | |||||||||||||||||||||||||
30 | 4/8/2026 | Diffusion LM | Diffusion-LM Improves Controllable Text Generation | |||||||||||||||||||||||
31 | Large Language Diffusion Models | |||||||||||||||||||||||||
32 | 4/13/2026 | Safety | On the Role of Attention Heads in Large Language Model Safety | |||||||||||||||||||||||
33 | Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs | |||||||||||||||||||||||||
34 | 4/15/2026 | Mechanistic Interpretability | Scaling and Evaluating Sparse Autoencoders | |||||||||||||||||||||||
35 | Sparse Crosscoders for Cross-Layer Features and Model Diffing | |||||||||||||||||||||||||
36 | 4/20/2026 | Calibration | Active Task Disambiguation with LLMs | |||||||||||||||||||||||
37 | Learning to Route LLMs with Confidence Tokens | |||||||||||||||||||||||||
38 | 4/22/2026 | Scaling Law | Training compute-optimal large language models | |||||||||||||||||||||||
39 | Scaling Laws for Precision | |||||||||||||||||||||||||
40 | ||||||||||||||||||||||||||
41 | ||||||||||||||||||||||||||
42 | ||||||||||||||||||||||||||
43 | ||||||||||||||||||||||||||
44 | ||||||||||||||||||||||||||
45 | ||||||||||||||||||||||||||
46 | ||||||||||||||||||||||||||
47 | ||||||||||||||||||||||||||
48 | ||||||||||||||||||||||||||
49 | ||||||||||||||||||||||||||
50 | ||||||||||||||||||||||||||
51 | ||||||||||||||||||||||||||
52 | ||||||||||||||||||||||||||
53 | ||||||||||||||||||||||||||
54 | ||||||||||||||||||||||||||
55 | ||||||||||||||||||||||||||
56 | ||||||||||||||||||||||||||
57 | ||||||||||||||||||||||||||
58 | ||||||||||||||||||||||||||
59 | ||||||||||||||||||||||||||
60 | ||||||||||||||||||||||||||
61 | ||||||||||||||||||||||||||
62 | ||||||||||||||||||||||||||
63 | ||||||||||||||||||||||||||
64 | ||||||||||||||||||||||||||
65 | ||||||||||||||||||||||||||
66 | ||||||||||||||||||||||||||
67 | ||||||||||||||||||||||||||
68 | ||||||||||||||||||||||||||
69 | ||||||||||||||||||||||||||
70 | ||||||||||||||||||||||||||
71 | ||||||||||||||||||||||||||
72 | ||||||||||||||||||||||||||
73 | ||||||||||||||||||||||||||
74 | ||||||||||||||||||||||||||
75 | ||||||||||||||||||||||||||
76 | ||||||||||||||||||||||||||
77 | ||||||||||||||||||||||||||
78 | ||||||||||||||||||||||||||
79 | ||||||||||||||||||||||||||
80 | ||||||||||||||||||||||||||
81 | ||||||||||||||||||||||||||
82 | ||||||||||||||||||||||||||
83 | ||||||||||||||||||||||||||
84 | ||||||||||||||||||||||||||
85 | ||||||||||||||||||||||||||
86 | ||||||||||||||||||||||||||
87 | ||||||||||||||||||||||||||
88 | ||||||||||||||||||||||||||
89 | ||||||||||||||||||||||||||
90 | ||||||||||||||||||||||||||
91 | ||||||||||||||||||||||||||
92 | ||||||||||||||||||||||||||
93 | ||||||||||||||||||||||||||
94 | ||||||||||||||||||||||||||
95 | ||||||||||||||||||||||||||
96 | ||||||||||||||||||||||||||
97 | ||||||||||||||||||||||||||
98 | ||||||||||||||||||||||||||
99 | ||||||||||||||||||||||||||
100 |