Good Papers

Showing Code generation Show all papers

83%Must read
?Must readVote to see the score

FrugalEvo: Towards Cost-Aware LLM-Guided Program Evolution

FrugalEvo pairs expensive LLM strategy exploration with cheap LLM implementation and caching to maximize optimization gain per cost under a budget, outperforming baselines on 10 tasks at significantly lower expense.

Hui Chen, Xuan Qi, James Zhao, Zhaopeng Feng and 4 more

Published Oct 2, 2026 · 0 citations · ▲ 22 on Hugging Face · Code ★ 5

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 1/5
89%Must read
?Must readVote to see the score

QuantCode Model: Specializing Language Models for Executable Algorithmic Trading Code

QuantCode specializes LLMs for executable trading code via framework pretraining and validated fine-tuning, boosting backtest success to 83.5% while revealing specialization trade-offs in tool use and repair.

Alexey Chernysh, Orkhan Ekhtibarov, Dmitry Zmitrovich

Published Sep 30, 2026 · ▲ 11 on Hugging Face

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 8/10
strict 3/5
78%Highly rated

Program-as-Weights: A Programming Paradigm for Fuzzy Functions

Program-as-Weights compiles natural-language specs into compact local neural adapters that match large-model prompting with far less memory and faster offline execution.

Wentao Zhang, Liliana Hotsko, Woojeong Kim, Pengyu Nie and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published Jul 2, 2026 · 0 citations · ▲ 307 on Hugging Face · Code ★ 359

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 5/10
strict 1/5
78%Highly rated
?Highly ratedVote to see the score

From Code Foundation Models to Agents and Applications: A Comprehensive Survey and Practical Guide to Code Intelligence

Survey synthesizes code LLM development from data curation to autonomous agents, analyzes general and specialized models, and maps the research-practice gap to practical needs.

Jian Yang, Xianglong Liu, Weifeng Lv, Ken Deng and 36 more

Published Nov 23, 2025 · 1 citation · ▲ 307 on Hugging Face

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 0/5
80%Must read
?Must readVote to see the score

On the Impacts of Contexts on Repository-Level Code Generation

RepoExec benchmarks repository-level code generation, showing instruction-tuned models use cross-file contexts better despite pretrained LLMs achieving higher correctness.

Nam Le Hai, Dung Manh Nguyen, Nghi D. Q. Bui

Published Jun 17, 2024 · 3 citations · ▲ 12 on Hugging Face · Code ★ 164

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
45%Niche pick
?Niche pickVote to see the score

SpecBridge: Learning Natural-Language Formalization Plans for the Formal Specification Synthesis Task

Wenjie Zhang, Yun Lin, Zining He, Meiqi Wu and 1 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

The Shape of a Program: Path Signatures for Trace-to-Program Induction

Mohamed Ghanem, Bernd Finkbeiner

Sydney Poster Session 1, Tue, Dec 8, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

Learning to Commit: Next-Commit Prediction via Online Supervised Contrastive Reflection

Mo Li, Qitai Tan, Kai Chen, Ting Cao and 1 more

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

The Locality Cost of Semantic Patch Self-Distillation

Xi Weng, Zhaoyu Liu, lianyu hu, Jin Song Dong

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

RepoZero: Can LLMs Generate a Code Repository from Scratch?

Zhaoxi Zhang, Yiming Xu, Jiahui Liang, Weikang Li and 6 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

ASAP: Assembly-Source Aligned Pseudocode Refinement For Binary Decompilation

Yujian Zhuang, Dehong Gao, Qichao Zhang, QiJing Lai and 3 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
67%Highly rated
?Highly ratedVote to see the score

ProgramBench: Can Language Models Rebuild Programs From Scratch?

John Yang, Kilian Lieret, Jeffrey Ma, Parth Thakkar and 8 more

Paris Poster Session 5, Fri, Dec 11, 11:30 AM–1:30 PM, Paris Poster Hall · Published 2026

– ReadersNo votes yet
2/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 2 of 20 reviewers recommend it
lenient 1/5
medium 1/10
strict 0/5
45%Niche pick
?Niche pickVote to see the score

How Do Language Models Compose Functions?

Apoorv Khandelwal, Ellie Pavlick

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
0/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 0 of 20 reviewers recommend it
lenient 0/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Phase-wise MLLM Tuning for Multi-framework WebUI Code Generation

Haoran Ma, Linxiao Li, Chenyue Wang, Haochen Sui and 1 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

Toward Executable Multi-framework Front-end Code Generation with Self-Correction

Linxiao Li, Haoran Ma, Chenyue Wang, Jiaye Lin and 2 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
57%Worth a look
?Worth a lookVote to see the score

HDL-RepoBench: Multi-Paradigm Repository-Level Code Completion for Hardware Design Languages

Qingyun Zou, Jiahao Cui, Nuo Chen, Bingsheng He and 1 more

Sydney Poster Session 3, Wed, Dec 9, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
1/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 1 of 20 reviewers recommend it
lenient 1/5
medium 0/10
strict 0/5
89%Must read
?Must readVote to see the score

ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization

ReLoop combines structured generation and behavioral verification to eliminate silent optimization formulation errors, reaching 100% executable code and improving accuracy across benchmarks.

Junbo Jacob Lian, Yujun Sun, Huiling Chen, Chaoyu Zhang and 2 more

Sydney Poster Session 2, Tue, Dec 8, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026 · ▲ 1 on Hugging Face · Code ★ 118

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 5/5
medium 9/10
strict 2/5
91%Must read
?Must readVote to see the score

P$^{3}$: Joint Program-and-Proof Planning\\ for Verified Code Generation

P³ plans programs and proofs jointly from specifications before elaboration, outperforming sequential baselines by up to 11.2 points on verified generation benchmarks while reducing cost and time.

Zenan Li, Ziran Yang, Peiyang Song, Zhaoyu Li and 1 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
18/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 18 of 20 reviewers recommend it
lenient 5/5
medium 10/10
strict 3/5
83%Must read
?Must readVote to see the score

Simple Baselines are Competitive with Code Evolution

Simple baselines match or beat sophisticated code evolution across mathematical bounds, agent scaffolds, and ML competitions, revealing evaluation flaws and underscoring that expert-designed search spaces matter more than search algorithms.

Yonatan Gideoni, Sebastian Risi, Yarin Gal

Sydney Poster Session 5, Thu, Dec 10, 10:00 AM–1:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
13/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 13 of 20 reviewers recommend it
lenient 5/5
medium 7/10
strict 1/5
80%Must read
?Must readVote to see the score

Natural Synthesis: Outperforming Reactive Synthesis Tools with Large Reasoning Models

A neuro-symbolic approach couples large reasoning models with model checkers to iteratively repair synthesized Verilog via sound symbolic feedback, solving more benchmarks than dedicated synthesis tools and enabling natural-language specification autoformalization.

Frederik Schmitt, Matthias Cosler, Niklas Metzger, Julian Siber and 3 more

Sydney Poster Session 6, Thu, Dec 10, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
12/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 12 of 20 reviewers recommend it
lenient 4/5
medium 6/10
strict 2/5
89%Must read
?Must readVote to see the score

Cross-Dialect Generalization Without Retraining: Benchmarks and Evaluation of Schema-Derived Constrained Decoding for MLIR

Schema-derived ODS constraints enable small LMs to match or exceed 15B, 34B models on structural MLIR dialects at 8, 25× speed without retraining, though attribute-heavy dialects remain challenging.

Plawan Kumar Rath

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
16/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 16 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 4/5
78%Highly rated
?Highly ratedVote to see the score

Constrained Code Generation with Discrete Diffusion

CDC integrates constraint satisfaction into discrete diffusion code generation via training-free neurosymbolic operators that locally steer denoising toward feasible programs, improving correctness, security, and syntax over baselines.

Lize Shao, Michael Cardei, Zichen Xie, Ferdinando Fioretto and 1 more

Atlanta Poster Session 4, Thu, Dec 10, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
11/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 11 of 20 reviewers recommend it
lenient 4/5
medium 7/10
strict 0/5
86%Must read
?Must readVote to see the score

From I/O to Code with Discovery Agent

DIO-Agent frames IO2Code as evolutionary search guided by execution errors and a simplicity-biased mutation prior, outperforming baselines on IO2CodeBench.

Yihong Dong, Jiaru Qian, Haoran Zhang, Peixu Wang and 6 more

Sydney Poster Session 4, Wed, Dec 9, 5:00 PM–8:00 PM, Hall 1-4 · Published 2026

– ReadersNo votes yet
14/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 14 of 20 reviewers recommend it
lenient 4/5
medium 8/10
strict 2/5
74%Highly rated
?Highly ratedVote to see the score

Viverra: Text-to-Code with Guarantees

Viverra generates C code with formally verified assertions via LLM synthesis and bounded model checking, improving user code comprehension.

Haoze Wu, Rocky Klopfenstein, Keith Farkas, Nina Narodytska

Atlanta Poster Session 6, Fri, Dec 11, 4:30 PM–7:30 PM, Hall C1 · Published 2026

– ReadersNo votes yet
9/20 AI panelreviewers recommend it

Readers and the AI panel: vote on this paper to see what they said.

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel: 9 of 20 reviewers recommend it
lenient 4/5
medium 4/10
strict 1/5