HuggingFace Papers
数据更新于 2026年09月08日 13:03(北京时间)
GitHub Trending
HF Trending
HF Papers
OpenAI
Anthropic
量子位
HuggingFace Papers
Today
Week
Month
共 105 条
1
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
2609.02749
▲ 529
2
StudentSim: Training LLM-based Student Simulators
2609.01591
▲ 483
3
Compile by Training: Turning Natural-Language Specifications into Local Neural Functions
2609.04199
▲ 376
4
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
2609.00111
▲ 376
5
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
2609.04148
▲ 281
6
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
2609.01437
▲ 256
7
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
2609.03796
▲ 228
8
Aspire: Can Models Self-Evolve from Vague Goals?
2608.31111
▲ 222
9
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
2609.03430
▲ 168
10
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training
2608.26730
▲ 149
11
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models
2609.02886
▲ 143
12
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
2608.31046
▲ 143
13
RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning
2609.03199
▲ 116
14
Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling
2608.30821
▲ 116
15
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction
2609.02783
▲ 115
16
LatentPress: Context Compression Beyond Text and Vision
2609.01507
▲ 111
17
Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
2609.02750
▲ 102
18
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
2609.01343
▲ 99
19
DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution
2608.31106
▲ 99
20
Rethinking On-Policy Distillation of Large Language Models II: One Training Example
2609.04172
▲ 82
21
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM
2609.04098
▲ 76
22
It Takes Two to Match: Co-Evolving Generative Retriever with Reinforcement Learning
2609.00638
▲ 72
23
Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States
2609.04196
▲ 66
24
GenFirst: Generation Before Reconstruction for Stable End-to-End Latent Generative Modeling
2608.29335
▲ 66
25
Language Models Can Control Their Own Attention
2609.02737
▲ 65
26
UI-Venus-2 Technical Report
2609.00028
▲ 62
27
On the Design of Qwen3.8-Next Architecture: Evaluation, Efficiency, and Training Stability
2608.30320
▲ 53
28
Iris: Climbing to the Search Frontier
2609.04304
▲ 51
29
ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training
2609.00188
▲ 51
30
H3-World: Turning Language Understanding into World Control
2609.01560
▲ 50
31
Normalized Low-Rank Adaptation
2608.31036
▲ 50
32
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction
2609.04201
▲ 46
33
Editable Visual Design
2609.04034
▲ 42
34
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
2608.31100
▲ 38
35
PaperGym: Rubric-Centered Evolution for Research-Plan Generation
2608.31119
▲ 38
36
Motion-Omni: End-to-End Joint Speech and Full-Body Motion for Spoken Dialogue
2609.04250
▲ 38
37
Hi-Q: Hierarchical Evidence-guided Query Refinement for Multi-Hop Question Answering
2608.30468
▲ 34
38
The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation
2609.02367
▲ 33
39
From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
2609.01572
▲ 33
40
On the Design Fundamentals of Pixel Text Representation Learning
2609.01147
▲ 32
41
LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation
2608.30935
▲ 32
42
WHALE: A Simple Recipe for Joint Harness-Weight Optimization
2609.00196
▲ 31
43
SHAPE of Chain-of-Thought in Math Reasoning
2608.28600
▲ 31
44
Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding
2609.04131
▲ 30
45
Last Translation Benchmark
2609.04173
▲ 30
46
Dr. Claw: An AI Scientist Workspace for Vibe Research
2609.00365
▲ 30
47
CogEvol: Towards Efficient and Reliable Learning Environment Generation
2608.30968
▲ 30
48
RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests
2608.27831
▲ 29
49
Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering
2608.21450
▲ 29
50
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence
2608.31075
▲ 28
51
Evaluating the Hidden Costs of Personalization in Large Language Models
2608.28833
▲ 28
52
Evaluating Multimodal LLMs as Generalist Vision-Language-Action Agents for Drone Control: Commanding, Approaching, Tracking and Searching
2609.01404
▲ 27
53
Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase
2608.29310
▲ 27
54
The Attention Triangle in Audio-Video Models
2609.03586
▲ 25
55
DRACO: Fine-Grained Credit Assignment with Dynamic Rubrics for Long-Horizon Agent Training
2609.04094
▲ 25
56
CORE: Improving Compositional Reasoning in MLLM Embedding via Reranker Distillation
2609.04083
▲ 25
57
ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes
2609.01740
▲ 25
58
NeoMME: A Single-Tower Multimodal-Native Multilingual Foundation Encoder for Efficient Fine-Tuning and Inference
2609.01657
▲ 25
59
WorldReward: Reward Modeling for Camera-Conditioned World Models
2609.03952
▲ 24
60
PACE: Towards Surfacing Hidden Conflicts in User Requests
2609.03293
▲ 23
61
Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System
2609.01607
▲ 23
62
WorldSculpt: Generating Compositional Worlds from Grounded Videos
2609.05416
▲ 21
63
DiagEvo: Diagnosis-Guided Self-Evolution via Hierarchical Error Memory
2609.00768
▲ 20
64
Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents
2608.31076
▲ 20
65
Safin-1: Safety from Within through Memory-Native State Evolution
2609.00092
▲ 20
66
FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow
2609.03563
▲ 19
67
Environment Evolution for Terminal Agents
2609.04128
▲ 19
68
Using Grounded Theory for Agent Behavior Analysis at Scale
2608.30391
▲ 19
69
Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
2608.29910
▲ 18
70
Cliff: Learning Process Rewards from the First Mistake
2609.02817
▲ 17
71
Enoki: Efficient Multi-Level Hallucination Detection
2609.00581
▲ 17
72
A Glance Is All You Need: Single-Pass Fine-Grained Image Captioning with SimLoss
2609.00591
▲ 17
73
Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
2609.01481
▲ 17
74
Principia: Relational Physics Tests for Video Models
2609.04200
▲ 16
75
VibeVoice-ASR-Streaming Technical Report
2609.02812
▲ 16
76
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
2608.30730
▲ 16
77
Influence-Directed Distillation: Solving the Diversity Bottleneck in Sampled-Token On-Policy Distillation
2608.29846
▲ 16
78
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
2609.05258
▲ 15
79
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
2609.05275
▲ 15
80
Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
2609.03820
▲ 15
81
Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance
2609.02373
▲ 14
82
EM^2Mem: Event-Centric Multimodal Memory for Large Language Models
2609.00551
▲ 14
83
Lies We Can See: Joint Verbal and Non-Verbal Deception by VLM Agents in Embodied Social Interactions
2608.30428
▲ 14
84
QCell: Recombining and Aligning Cell Queries for Overlapping Instance Segmentation
2608.29253
▲ 14
85
UniMate: One Unified Model to Animate Diverse Skeletons
2609.05415
▲ 13
86
Let Confidence Change, Not the Prediction: Prediction-Preserving Repair for Post-hoc Calibration
2609.01072
▲ 13
87
Chain-of-Thought Faithfulness of Reasoning Models Varies with Where and How Preference Cues Are Delivered
2608.29464
▲ 13
88
Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation
2608.24293
▲ 13
89
VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement
2609.03153
▲ 12
90
PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback
2608.30241
▲ 12
91
Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers
2608.18972
▲ 11
92
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
2609.04753
▲ 10
93
MaxKernel: Agentic Kernel Generation for TPUs
2609.04523
▲ 10
94
A Common Measure of Communication for Speech Brain-Computer Interfaces
2609.02887
▲ 10
95
Agents in the Large: Perception-Centered Architecture for Persistent Agents
2608.30478
▲ 10
96
Verification-Aware Training for Speculative Decoding
2608.30135
▲ 10
97
Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation
2608.30396
▲ 10
98
Locked at the Entrance, Open Inside: Where RLVR Narrows the Solution Space
2608.29188
▲ 10
99
RISE: Recursive Improvement via Self-Extrapolating Policy Distillation
2609.05295
▲ 9
100
Post-Training Language Models for Gold-Medal Performance in Coding Competitions
2609.02849
▲ 9
101
Learning Where Outcomes Change:Credit-Addressable Reasoning for Multimodal Geometry
2608.30457
▲ 9
102
CRISP: Cliff-awaRe Input-adaptive Sparse Prefilling with Structural-Mass-Motivated Routing
2609.01925
▲ 8
103
Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall
2609.01532
▲ 8
104
MULTI3IR: A Benchmark for Multi-perspective Multi-domain Multi-modal Information Retrieval
2608.30949
▲ 8
105
WebWorld: The Browser as a World Model for Self-Improving Web Code
2608.30530
▲ 8
没有匹配项