Commit0: Library Generation from Scratch

Na minha lista:
Detalhes bibliográficos
Publicado no:arXiv.org (Dec 2, 2024), p. n/a
Autor principal: Zhao, Wenting
Outros Autores: Jiang, Nan, Lee, Celine, Chiu, Justin T, Cardie, Claire, Gallé, Matthias, Rush, Alexander M
Publicado em:
Cornell University Library, arXiv.org
Assuntos:
Acesso em linha:Citation/Abstract
Full text outside of ProQuest
Tags: Adicionar Tag
Sem tags, seja o primeiro a adicionar uma tag!

MARC

LEADER 00000nab a2200000uu 4500
001 3138994667
003 UK-CbPIL
022 |a 2331-8422 
035 |a 3138994667 
045 0 |b d20241202 
100 1 |a Zhao, Wenting 
245 1 |a Commit0: Library Generation from Scratch 
260 |b Cornell University Library, arXiv.org  |c Dec 2, 2024 
513 |a Working Paper 
520 3 |a With the goal of benchmarking generative systems beyond expert software development ability, we introduce Commit0, a benchmark that challenges AI agents to write libraries from scratch. Agents are provided with a specification document outlining the library's API as well as a suite of interactive unit tests, with the goal of producing an implementation of this API accordingly. The implementation is validated through running these unit tests. As a benchmark, Commit0 is designed to move beyond static one-shot code generation towards agents that must process long-form natural language specifications, adapt to multi-stage feedback, and generate code with complex dependencies. Commit0 also offers an interactive environment where models receive static analysis and execution feedback on the code they generate. Our experiments demonstrate that while current agents can pass some unit tests, none can yet fully reproduce full libraries. Results also show that interactive feedback is quite useful for models to generate code that passes more unit tests, validating the benchmarks that facilitate its use. 
653 |a Application programming interface 
653 |a Static code analysis 
653 |a Feedback 
653 |a Specifications 
653 |a Software development 
653 |a Benchmarks 
653 |a Speech recognition 
700 1 |a Jiang, Nan 
700 1 |a Lee, Celine 
700 1 |a Chiu, Justin T 
700 1 |a Cardie, Claire 
700 1 |a Gallé, Matthias 
700 1 |a Rush, Alexander M 
773 0 |t arXiv.org  |g (Dec 2, 2024), p. n/a 
786 0 |d ProQuest  |t Engineering Database 
856 4 1 |3 Citation/Abstract  |u https://www.proquest.com/docview/3138994667/abstract/embedded/ZKJTFFSVAI7CB62C?source=fedsrch 
856 4 0 |3 Full text outside of ProQuest  |u http://arxiv.org/abs/2412.01769