ABCDEFGHIJKLMNOPQRSTUVWXYZ
1
DateTopicPapers
2
2/2/2026PretrainingParity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
3
Chameleon: A Flexible Data-mixing Framework for Language Model Pretraining and Finetuning
4
2/4/2026EmbeddingsMMTEB: Massive Multilingual Text Embedding Benchmark
5
Improving Text Embeddings with Large Language Models
6
2/9/2026EmbeddingsDeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model (only Section 2.1)
7
NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models
8
2/11/2026Mixture-of-ExpertsOutrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
9
Your Mixture-of-Experts LLM Is Secretly an Embedding Model for Free
10
2/16/2026ReasoningBIRD: A Trustworthy Bayesian Inference Framework for Large Language Models
11
2/18/2026Reasoning
/ Alignment
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
12
Training a Generally Curious Agent
13
2/23/2026Agent HarnessAgent Workflow Memory
14
The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents
15
2/25/2026Long-ContextEfficient Streaming Language Models with Attention Sinks
16
3/2/2026Latent AttentionTransMLA: Multi-Head Latent Attention Is All You Need
17
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models (only Section 2)
18
3/4/2026Reinforcement LearningDAPO: An Open-Source LLM Reinforcement Learning System at Scale
19
Understanding R1-Zero-Like Training
20
3/9/2026ReasoningOptimas: Optimizing Compound AI Systems with Globally Aligned Local Rewards
21
3/11/2026ReasoningCWM: An Open-Weights LLM for Research on Code Generation with World Models
22
3/16/2026Self-Play RLAbsolute Zero: Reinforced Self-play Reasoning with Zero Data
23
3/18/2026Self-Play RLSPICE: Self-Play In Corpus Environments Improves Reasoning
24
Toward Training Superintelligent Software Agents through Self-Play SWE-RL
25
3/30/2026Test-Time ScalingScaling LLM Test-Time Compute Optimally Can be More Effective than Scaling Parameters for Reasoning
26
Learning to Discover at Test Time
27
4/1/2026Mode CollapseArtificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
28
4/6/2026Linear TransformerLinear transformers are secretly fast weight programmers
29
Parallelizing Linear Transformers with the Delta Rule over Sequence Length
30
4/8/2026Diffusion LMDiffusion-LM Improves Controllable Text Generation
31
Large Language Diffusion Models
32
4/13/2026SafetyOn the Role of Attention Heads in Large Language Model Safety
33
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
34
4/15/2026Mechanistic InterpretabilityScaling and Evaluating Sparse Autoencoders
35
Sparse Crosscoders for Cross-Layer Features and Model Diffing
36
4/20/2026CalibrationActive Task Disambiguation with LLMs
37
Learning to Route LLMs with Confidence Tokens
38
4/22/2026Scaling LawTraining compute-optimal large language models
39
Scaling Laws for Precision
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100