Good Papers

On the Impacts of Contexts on Repository-Level Code Generation

RepoExec benchmarks repository-level code generation, showing instruction-tuned models use cross-file contexts better despite pretrained LLMs achieving higher correctness.

Nam Le Hai, Dung Manh Nguyen, Nghi D. Q. Bui

Published Jun 17, 20243 citations▲ 12 on Hugging FaceCode ★ 164arXiv ↗

80%
OverallMust read
?
OverallMust readVote to see the scoreThe exact score shows once you've voted, so every vote is your own call. The first half of each home page shelf shows its scores.
Readers
–

Only vote on papers you've read. Sign in with GitHub to vote.

AI panel12/20reviewers recommend it
lenient 5/5
medium 6/10
strict 1/5
AI panel?Vote to see what the 20 AI reviewers said
Panel consensus
RepoExec is a vital, timely benchmark that isolates cross-file context utilization with a useful DIR metric and instruction-tuning split, though its synthetic dependency specs and missing multi-file edit baselines leave real-world integration quality uncertain.

Abstract

CodeLLMs have gained widespread adoption for code generation tasks, yet their capacity to handle repository-level code generation with complex contextual dependencies remains underexplored. Our work underscores the critical importance of leveraging repository-level contexts to generate executable and functionally correct code. We present RepoExec, a novel benchmark designed to evaluate repository-level code generation, with a focus on three key aspects: executability, functional correctness through comprehensive test case generation, and accurate utilization of cross-file contexts. Our study examines a controlled scenario where developers specify essential code dependencies (contexts), challenging models to integrate them effectively. Additionally, we introduce an instruction-tuned dataset that enhances CodeLLMs' ability to leverage dependencies, along with a new metric, Dependency Invocation Rate (DIR), to quantify context utilization. Experimental results reveal that while pretrained LLMs demonstrate superior performance in terms of correctness, instruction-tuned models excel in context utilization and debugging capabilities. RepoExec offers a comprehensive evaluation framework for assessing code functionality and alignment with developer intent, thereby advancing the development of more reliable CodeLLMs for real-world applications. The dataset and source code are available at https://github.com/FSoft-AI4Code/RepoExec.