Skip to content

Paper Catalog

This page is generated from the reviewed database. The grouped sections support theme browsing, and the full registry below is optimized for MkDocs search.

Grouped By Review Dimension

RQ1 / Knowledge Layers

Repository Knowledge (128)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026empowering (2026): Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yu2026does (2026): Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chakraborty2025blaze (2025): BLAZE: Cross-language and cross-project bug localization via dynamic chunking and hard example learning

  • chang2025bridging (2025): Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models

  • chen2025locagent (2025): Locagent: Graph-guided llm agents for code localization

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • dai2025lita (2025): Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

  • ehrlich2025codemonkeys (2025): Codemonkeys: Scaling test-time compute for software engineering

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • gupta2025sacl (2025): SACL: Understanding and Combating Textual Bias in Code Retrieval with Semantic-Augmented Reranking and Localization

  • huang2025seeing (2025): Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • jiang2025cosil (2025): CoSIL: Software Issue Localization via LLM-Driven Code Repository Graph Searching

  • jiang2025putting (2025): Putting It All into Context: Simplifying Agents with LCLMs

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025infcode (2025): InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • liu2025graphlocator (2025): GraphLocator: Graph-guided Causal Reasoning for Issue Localization

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025sorft (2025): Sorft: Issue resolving with subtask-oriented reinforced fine-tuning

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • pabba2025semagent (2025): SemAgent: A Semantics Aware Program Repair Agent

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • reddy2025swerank (2025): SweRank: Software Issue Localization with Code Ranking

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • ruan2025specrover (2025): Specrover: Code intent extraction via llms

  • shi2025aime (2025): Aime: Towards Fully-Autonomous Multi-Agent Framework

  • shum2025swe (2025): SWE-RM: Execution-free Feedback For Software Engineering Agents

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • soni2025coding (2025): Coding Agents with Multimodal Browsing are Generalist Problem Solvers

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • tao2025code (2025): Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • tawosi2025meta (2025): Meta-RAG on Large Codebases Using Code Summarization

  • vaghasiya2025corethink (2025): CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs

  • vinh2025repeton (2025): Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wang2025extracting (2025): Extracting Conceptual Knowledge to Locate Software Issues

  • wang2025improving (2025): Improving Code Localization with Repository Memory

  • wang2025mcts (2025): Mcts-refined cot: High-quality fine-tuning data for llm-based repository issue resolution

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025swe (2025): Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • xiong2025think (2025): Think-Search-Patch: A Retrieval-Augmented Reasoning Framework for Repository-Level Code Repair

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • yu2025utboost (2025): Utboost: Rigorous evaluation of coding agents on swe-bench

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025cast (2025): cast: Enhancing code retrieval-augmented generation with structural chunking via abstract syntax tree

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • arora2024masai (2024): Masai: Modular architecture for software-engineering ai agents

  • brown2024large (2024): Large language monkeys: Scaling inference compute with repeated sampling

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • cheshkov2024exploring (2024): Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • li2024codetree (2024): Codetree: Agent-guided tree search for code generation with large language models

  • lin2024llms (2024): Llms as continuous learners: Improving the reproduction of defective code in software issues

  • liu2024codexembed (2024): Codexembed: A generalist embedding model family for multiligual and multi-task code retrieval

  • liu2024codexgraph (2024): Codexgraph: Bridging large language models and code repositories via code graph databases

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • ma2024repository (2024): Repository Structure-Aware Training Makes SLMs Better Issue Resolver

  • ouyang2024repograph (2024): Repograph: Enhancing ai software engineering with repository-level code graph

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • rombaut2024watson (2024): Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents

  • suresh2024cornstack (2024): CoRNStack: High-quality contrastive data for better code retrieval and reranking

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024autocoderover (2024): Autocoderover: Autonomous program improvement

  • zhang2024codev (2024): CodeV: Issue Resolving with Visual Data

  • zhang2024diversity (2024): Diversity empowers intelligence: Integrating expertise of software engineering agents

Procedural Knowledge (109)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026empowering (2026): Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chen2025locagent (2025): Locagent: Graph-guided llm agents for code localization

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • dai2025lita (2025): Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

  • ehrlich2025codemonkeys (2025): Codemonkeys: Scaling test-time compute for software engineering

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • jiang2025putting (2025): Putting It All into Context: Simplifying Agents with LCLMs

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025infcode (2025): InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • liu2025graphlocator (2025): GraphLocator: Graph-guided Causal Reasoning for Issue Localization

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • pabba2025semagent (2025): SemAgent: A Semantics Aware Program Repair Agent

  • pan2025codecor (2025): Codecor: An llm-based self-reflective multi-agent framework for code generation

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • ruan2025specrover (2025): Specrover: Code intent extraction via llms

  • shi2025aime (2025): Aime: Towards Fully-Autonomous Multi-Agent Framework

  • shum2025swe (2025): SWE-RM: Execution-free Feedback For Software Engineering Agents

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • soni2025coding (2025): Coding Agents with Multimodal Browsing are Generalist Problem Solvers

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • su2025learn (2025): Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025agent (2025): Agent kb: Leveraging cross-domain experience for agentic problem solving

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • vaghasiya2025corethink (2025): CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs

  • vinh2025repeton (2025): Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • yu2025utboost (2025): Utboost: Rigorous evaluation of coding agents on swe-bench

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhang2025sealign (2025): Sealign: Alignment training for software engineering agent

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • arora2024masai (2024): Masai: Modular architecture for software-engineering ai agents

  • brown2024large (2024): Large language monkeys: Scaling inference compute with repeated sampling

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • cheshkov2024exploring (2024): Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • lin2024llms (2024): Llms as continuous learners: Improving the reproduction of defective code in software issues

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • rombaut2024watson (2024): Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024autocoderover (2024): Autocoderover: Autonomous program improvement

  • zhang2024diversity (2024): Diversity empowers intelligence: Integrating expertise of software engineering agents

Background Knowledge (6)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

RQ1 / Knowledge Types

Existing Code Knowledge (128)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026empowering (2026): Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yu2026does (2026): Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chakraborty2025blaze (2025): BLAZE: Cross-language and cross-project bug localization via dynamic chunking and hard example learning

  • chang2025bridging (2025): Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models

  • chen2025locagent (2025): Locagent: Graph-guided llm agents for code localization

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • dai2025lita (2025): Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

  • ehrlich2025codemonkeys (2025): Codemonkeys: Scaling test-time compute for software engineering

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • gupta2025sacl (2025): SACL: Understanding and Combating Textual Bias in Code Retrieval with Semantic-Augmented Reranking and Localization

  • huang2025seeing (2025): Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • jiang2025cosil (2025): CoSIL: Software Issue Localization via LLM-Driven Code Repository Graph Searching

  • jiang2025putting (2025): Putting It All into Context: Simplifying Agents with LCLMs

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025infcode (2025): InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • liu2025graphlocator (2025): GraphLocator: Graph-guided Causal Reasoning for Issue Localization

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025sorft (2025): Sorft: Issue resolving with subtask-oriented reinforced fine-tuning

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • pabba2025semagent (2025): SemAgent: A Semantics Aware Program Repair Agent

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • reddy2025swerank (2025): SweRank: Software Issue Localization with Code Ranking

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • ruan2025specrover (2025): Specrover: Code intent extraction via llms

  • shi2025aime (2025): Aime: Towards Fully-Autonomous Multi-Agent Framework

  • shum2025swe (2025): SWE-RM: Execution-free Feedback For Software Engineering Agents

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • soni2025coding (2025): Coding Agents with Multimodal Browsing are Generalist Problem Solvers

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • tao2025code (2025): Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • tawosi2025meta (2025): Meta-RAG on Large Codebases Using Code Summarization

  • vaghasiya2025corethink (2025): CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs

  • vinh2025repeton (2025): Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wang2025extracting (2025): Extracting Conceptual Knowledge to Locate Software Issues

  • wang2025improving (2025): Improving Code Localization with Repository Memory

  • wang2025mcts (2025): Mcts-refined cot: High-quality fine-tuning data for llm-based repository issue resolution

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025swe (2025): Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • xiong2025think (2025): Think-Search-Patch: A Retrieval-Augmented Reasoning Framework for Repository-Level Code Repair

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • yu2025utboost (2025): Utboost: Rigorous evaluation of coding agents on swe-bench

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025cast (2025): cast: Enhancing code retrieval-augmented generation with structural chunking via abstract syntax tree

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • arora2024masai (2024): Masai: Modular architecture for software-engineering ai agents

  • brown2024large (2024): Large language monkeys: Scaling inference compute with repeated sampling

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • cheshkov2024exploring (2024): Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • li2024codetree (2024): Codetree: Agent-guided tree search for code generation with large language models

  • lin2024llms (2024): Llms as continuous learners: Improving the reproduction of defective code in software issues

  • liu2024codexembed (2024): Codexembed: A generalist embedding model family for multiligual and multi-task code retrieval

  • liu2024codexgraph (2024): Codexgraph: Bridging large language models and code repositories via code graph databases

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • ma2024repository (2024): Repository Structure-Aware Training Makes SLMs Better Issue Resolver

  • ouyang2024repograph (2024): Repograph: Enhancing ai software engineering with repository-level code graph

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • rombaut2024watson (2024): Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents

  • suresh2024cornstack (2024): CoRNStack: High-quality contrastive data for better code retrieval and reranking

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024autocoderover (2024): Autocoderover: Autonomous program improvement

  • zhang2024codev (2024): CodeV: Issue Resolving with Visual Data

  • zhang2024diversity (2024): Diversity empowers intelligence: Integrating expertise of software engineering agents

Development Tool Knowledge (103)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026empowering (2026): Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • dai2025lita (2025): Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

  • ehrlich2025codemonkeys (2025): Codemonkeys: Scaling test-time compute for software engineering

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • jiang2025putting (2025): Putting It All into Context: Simplifying Agents with LCLMs

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025infcode (2025): InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • liu2025graphlocator (2025): GraphLocator: Graph-guided Causal Reasoning for Issue Localization

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • pabba2025semagent (2025): SemAgent: A Semantics Aware Program Repair Agent

  • pan2025codecor (2025): Codecor: An llm-based self-reflective multi-agent framework for code generation

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • ruan2025specrover (2025): Specrover: Code intent extraction via llms

  • shi2025aime (2025): Aime: Towards Fully-Autonomous Multi-Agent Framework

  • shum2025swe (2025): SWE-RM: Execution-free Feedback For Software Engineering Agents

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • soni2025coding (2025): Coding Agents with Multimodal Browsing are Generalist Problem Solvers

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • vaghasiya2025corethink (2025): CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs

  • vinh2025repeton (2025): Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • yu2025utboost (2025): Utboost: Rigorous evaluation of coding agents on swe-bench

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhang2025sealign (2025): Sealign: Alignment training for software engineering agent

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • arora2024masai (2024): Masai: Modular architecture for software-engineering ai agents

  • brown2024large (2024): Large language monkeys: Scaling inference compute with repeated sampling

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • cheshkov2024exploring (2024): Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • lin2024llms (2024): Llms as continuous learners: Improving the reproduction of defective code in software issues

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024autocoderover (2024): Autocoderover: Autonomous program improvement

Experiential Knowledge (68)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chen2025locagent (2025): Locagent: Graph-guided llm agents for code localization

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • shum2025swe (2025): SWE-RM: Execution-free Feedback For Software Engineering Agents

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • su2025learn (2025): Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments

  • tang2025agent (2025): Agent kb: Leveraging cross-domain experience for agentic problem solving

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhang2025sealign (2025): Sealign: Alignment training for software engineering agent

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • lin2024llms (2024): Llms as continuous learners: Improving the reproduction of defective code in software issues

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • rombaut2024watson (2024): Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024diversity (2024): Diversity empowers intelligence: Integrating expertise of software engineering agents

Repository Evolution Knowledge (43)

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yu2026does (2026): Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chakraborty2025blaze (2025): BLAZE: Cross-language and cross-project bug localization via dynamic chunking and hard example learning

  • chang2025bridging (2025): Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • ma2025sorft (2025): Sorft: Issue resolving with subtask-oriented reinforced fine-tuning

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • reddy2025swerank (2025): SweRank: Software Issue Localization with Code Ranking

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tao2025code (2025): Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • wang2025improving (2025): Improving Code Localization with Repository Memory

  • wang2025mcts (2025): Mcts-refined cot: High-quality fine-tuning data for llm-based repository issue resolution

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025swe (2025): Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025utboost (2025): Utboost: Rigorous evaluation of coding agents on swe-bench

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • liu2024codexembed (2024): Codexembed: A generalist embedding model family for multiligual and multi-task code retrieval

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • ma2024repository (2024): Repository Structure-Aware Training Makes SLMs Better Issue Resolver

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

Design Architecture Knowledge (13)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • yu2026does (2026): Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • huang2025seeing (2025): Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

Technology Stack Knowledge (4)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

Programming Language Knowledge (3)

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

Domain-specific Knowledge (1)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

RQ2 / Knowledge Sources

Codebase (128)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026empowering (2026): Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yu2026does (2026): Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chakraborty2025blaze (2025): BLAZE: Cross-language and cross-project bug localization via dynamic chunking and hard example learning

  • chang2025bridging (2025): Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models

  • chen2025locagent (2025): Locagent: Graph-guided llm agents for code localization

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • dai2025lita (2025): Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

  • ehrlich2025codemonkeys (2025): Codemonkeys: Scaling test-time compute for software engineering

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • gupta2025sacl (2025): SACL: Understanding and Combating Textual Bias in Code Retrieval with Semantic-Augmented Reranking and Localization

  • huang2025seeing (2025): Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • jiang2025cosil (2025): CoSIL: Software Issue Localization via LLM-Driven Code Repository Graph Searching

  • jiang2025putting (2025): Putting It All into Context: Simplifying Agents with LCLMs

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025infcode (2025): InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • liu2025graphlocator (2025): GraphLocator: Graph-guided Causal Reasoning for Issue Localization

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025sorft (2025): Sorft: Issue resolving with subtask-oriented reinforced fine-tuning

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • pabba2025semagent (2025): SemAgent: A Semantics Aware Program Repair Agent

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • reddy2025swerank (2025): SweRank: Software Issue Localization with Code Ranking

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • ruan2025specrover (2025): Specrover: Code intent extraction via llms

  • shi2025aime (2025): Aime: Towards Fully-Autonomous Multi-Agent Framework

  • shum2025swe (2025): SWE-RM: Execution-free Feedback For Software Engineering Agents

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • soni2025coding (2025): Coding Agents with Multimodal Browsing are Generalist Problem Solvers

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • tao2025code (2025): Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • tawosi2025meta (2025): Meta-RAG on Large Codebases Using Code Summarization

  • vaghasiya2025corethink (2025): CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs

  • vinh2025repeton (2025): Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wang2025extracting (2025): Extracting Conceptual Knowledge to Locate Software Issues

  • wang2025improving (2025): Improving Code Localization with Repository Memory

  • wang2025mcts (2025): Mcts-refined cot: High-quality fine-tuning data for llm-based repository issue resolution

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025swe (2025): Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • xiong2025think (2025): Think-Search-Patch: A Retrieval-Augmented Reasoning Framework for Repository-Level Code Repair

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • yu2025utboost (2025): Utboost: Rigorous evaluation of coding agents on swe-bench

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025cast (2025): cast: Enhancing code retrieval-augmented generation with structural chunking via abstract syntax tree

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • arora2024masai (2024): Masai: Modular architecture for software-engineering ai agents

  • brown2024large (2024): Large language monkeys: Scaling inference compute with repeated sampling

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • cheshkov2024exploring (2024): Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • li2024codetree (2024): Codetree: Agent-guided tree search for code generation with large language models

  • lin2024llms (2024): Llms as continuous learners: Improving the reproduction of defective code in software issues

  • liu2024codexembed (2024): Codexembed: A generalist embedding model family for multiligual and multi-task code retrieval

  • liu2024codexgraph (2024): Codexgraph: Bridging large language models and code repositories via code graph databases

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • ma2024repository (2024): Repository Structure-Aware Training Makes SLMs Better Issue Resolver

  • ouyang2024repograph (2024): Repograph: Enhancing ai software engineering with repository-level code graph

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • rombaut2024watson (2024): Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents

  • suresh2024cornstack (2024): CoRNStack: High-quality contrastive data for better code retrieval and reranking

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024autocoderover (2024): Autocoderover: Autonomous program improvement

  • zhang2024codev (2024): CodeV: Issue Resolving with Visual Data

  • zhang2024diversity (2024): Diversity empowers intelligence: Integrating expertise of software engineering agents

Dynamic Execution Results (105)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026empowering (2026): Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chen2025locagent (2025): Locagent: Graph-guided llm agents for code localization

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • dai2025lita (2025): Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

  • ehrlich2025codemonkeys (2025): Codemonkeys: Scaling test-time compute for software engineering

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • jiang2025putting (2025): Putting It All into Context: Simplifying Agents with LCLMs

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025infcode (2025): InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • pabba2025semagent (2025): SemAgent: A Semantics Aware Program Repair Agent

  • pan2025codecor (2025): Codecor: An llm-based self-reflective multi-agent framework for code generation

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • ruan2025specrover (2025): Specrover: Code intent extraction via llms

  • shi2025aime (2025): Aime: Towards Fully-Autonomous Multi-Agent Framework

  • shum2025swe (2025): SWE-RM: Execution-free Feedback For Software Engineering Agents

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • soni2025coding (2025): Coding Agents with Multimodal Browsing are Generalist Problem Solvers

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • su2025learn (2025): Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025agent (2025): Agent kb: Leveraging cross-domain experience for agentic problem solving

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • vaghasiya2025corethink (2025): CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs

  • vinh2025repeton (2025): Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • yu2025utboost (2025): Utboost: Rigorous evaluation of coding agents on swe-bench

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhang2025sealign (2025): Sealign: Alignment training for software engineering agent

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • arora2024masai (2024): Masai: Modular architecture for software-engineering ai agents

  • brown2024large (2024): Large language monkeys: Scaling inference compute with repeated sampling

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • cheshkov2024exploring (2024): Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • lin2024llms (2024): Llms as continuous learners: Improving the reproduction of defective code in software issues

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • rombaut2024watson (2024): Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024autocoderover (2024): Autocoderover: Autonomous program improvement

Expert (43)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026empowering (2026): Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • liu2025graphlocator (2025): GraphLocator: Graph-guided Causal Reasoning for Issue Localization

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • tang2025agent (2025): Agent kb: Leveraging cross-domain experience for agentic problem solving

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024diversity (2024): Diversity empowers intelligence: Integrating expertise of software engineering agents

Repository History (43)

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yu2026does (2026): Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chakraborty2025blaze (2025): BLAZE: Cross-language and cross-project bug localization via dynamic chunking and hard example learning

  • chang2025bridging (2025): Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • ma2025sorft (2025): Sorft: Issue resolving with subtask-oriented reinforced fine-tuning

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • reddy2025swerank (2025): SweRank: Software Issue Localization with Code Ranking

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tao2025code (2025): Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • wang2025improving (2025): Improving Code Localization with Repository Memory

  • wang2025mcts (2025): Mcts-refined cot: High-quality fine-tuning data for llm-based repository issue resolution

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025swe (2025): Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025utboost (2025): Utboost: Rigorous evaluation of coding agents on swe-bench

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • liu2024codexembed (2024): Codexembed: A generalist embedding model family for multiligual and multi-task code retrieval

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • ma2024repository (2024): Repository Structure-Aware Training Makes SLMs Better Issue Resolver

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

Document (11)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • huang2025seeing (2025): Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing

  • su2025learn (2025): Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

RQ2 / Extraction Methods

Rule-based Extraction (127)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026empowering (2026): Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yu2026does (2026): Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chakraborty2025blaze (2025): BLAZE: Cross-language and cross-project bug localization via dynamic chunking and hard example learning

  • chang2025bridging (2025): Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models

  • chen2025locagent (2025): Locagent: Graph-guided llm agents for code localization

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • dai2025lita (2025): Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

  • ehrlich2025codemonkeys (2025): Codemonkeys: Scaling test-time compute for software engineering

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • gupta2025sacl (2025): SACL: Understanding and Combating Textual Bias in Code Retrieval with Semantic-Augmented Reranking and Localization

  • huang2025seeing (2025): Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • jiang2025cosil (2025): CoSIL: Software Issue Localization via LLM-Driven Code Repository Graph Searching

  • jiang2025putting (2025): Putting It All into Context: Simplifying Agents with LCLMs

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025infcode (2025): InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • liu2025graphlocator (2025): GraphLocator: Graph-guided Causal Reasoning for Issue Localization

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025sorft (2025): Sorft: Issue resolving with subtask-oriented reinforced fine-tuning

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • pabba2025semagent (2025): SemAgent: A Semantics Aware Program Repair Agent

  • pan2025codecor (2025): Codecor: An llm-based self-reflective multi-agent framework for code generation

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • reddy2025swerank (2025): SweRank: Software Issue Localization with Code Ranking

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • ruan2025specrover (2025): Specrover: Code intent extraction via llms

  • shi2025aime (2025): Aime: Towards Fully-Autonomous Multi-Agent Framework

  • shum2025swe (2025): SWE-RM: Execution-free Feedback For Software Engineering Agents

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • soni2025coding (2025): Coding Agents with Multimodal Browsing are Generalist Problem Solvers

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • tao2025code (2025): Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • tawosi2025meta (2025): Meta-RAG on Large Codebases Using Code Summarization

  • vaghasiya2025corethink (2025): CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs

  • vinh2025repeton (2025): Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wang2025extracting (2025): Extracting Conceptual Knowledge to Locate Software Issues

  • wang2025improving (2025): Improving Code Localization with Repository Memory

  • wang2025mcts (2025): Mcts-refined cot: High-quality fine-tuning data for llm-based repository issue resolution

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025swe (2025): Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • xiong2025think (2025): Think-Search-Patch: A Retrieval-Augmented Reasoning Framework for Repository-Level Code Repair

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • yu2025utboost (2025): Utboost: Rigorous evaluation of coding agents on swe-bench

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025cast (2025): cast: Enhancing code retrieval-augmented generation with structural chunking via abstract syntax tree

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhang2025sealign (2025): Sealign: Alignment training for software engineering agent

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • arora2024masai (2024): Masai: Modular architecture for software-engineering ai agents

  • brown2024large (2024): Large language monkeys: Scaling inference compute with repeated sampling

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • cheshkov2024exploring (2024): Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • lin2024llms (2024): Llms as continuous learners: Improving the reproduction of defective code in software issues

  • liu2024codexembed (2024): Codexembed: A generalist embedding model family for multiligual and multi-task code retrieval

  • liu2024codexgraph (2024): Codexgraph: Bridging large language models and code repositories via code graph databases

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • ma2024repository (2024): Repository Structure-Aware Training Makes SLMs Better Issue Resolver

  • ouyang2024repograph (2024): Repograph: Enhancing ai software engineering with repository-level code graph

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • suresh2024cornstack (2024): CoRNStack: High-quality contrastive data for better code retrieval and reranking

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024autocoderover (2024): Autocoderover: Autonomous program improvement

  • zhang2024codev (2024): CodeV: Issue Resolving with Visual Data

  • zhang2024diversity (2024): Diversity empowers intelligence: Integrating expertise of software engineering agents

Model-based Extraction (91)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yu2026does (2026): Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • chang2025bridging (2025): Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models

  • chen2025locagent (2025): Locagent: Graph-guided llm agents for code localization

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • ehrlich2025codemonkeys (2025): Codemonkeys: Scaling test-time compute for software engineering

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • gupta2025sacl (2025): SACL: Understanding and Combating Textual Bias in Code Retrieval with Semantic-Augmented Reranking and Localization

  • huang2025seeing (2025): Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • jiang2025cosil (2025): CoSIL: Software Issue Localization via LLM-Driven Code Repository Graph Searching

  • jiang2025putting (2025): Putting It All into Context: Simplifying Agents with LCLMs

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • pabba2025semagent (2025): SemAgent: A Semantics Aware Program Repair Agent

  • reddy2025swerank (2025): SweRank: Software Issue Localization with Code Ranking

  • ruan2025specrover (2025): Specrover: Code intent extraction via llms

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • su2025learn (2025): Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments

  • tang2025agent (2025): Agent kb: Leveraging cross-domain experience for agentic problem solving

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • tao2025code (2025): Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • tawosi2025meta (2025): Meta-RAG on Large Codebases Using Code Summarization

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wang2025extracting (2025): Extracting Conceptual Knowledge to Locate Software Issues

  • wang2025improving (2025): Improving Code Localization with Repository Memory

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025swe (2025): Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • xiong2025think (2025): Think-Search-Patch: A Retrieval-Augmented Reasoning Framework for Repository-Level Code Repair

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhang2025sealign (2025): Sealign: Alignment training for software engineering agent

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • arora2024masai (2024): Masai: Modular architecture for software-engineering ai agents

  • brown2024large (2024): Large language monkeys: Scaling inference compute with repeated sampling

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • li2024codetree (2024): Codetree: Agent-guided tree search for code generation with large language models

  • lin2024llms (2024): Llms as continuous learners: Improving the reproduction of defective code in software issues

  • liu2024codexembed (2024): Codexembed: A generalist embedding model family for multiligual and multi-task code retrieval

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • ouyang2024repograph (2024): Repograph: Enhancing ai software engineering with repository-level code graph

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • rombaut2024watson (2024): Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

  • zhang2024diversity (2024): Diversity empowers intelligence: Integrating expertise of software engineering agents

Manual Extraction (39)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026empowering (2026): Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • liu2025graphlocator (2025): GraphLocator: Graph-guided Causal Reasoning for Issue Localization

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025agent (2025): Agent kb: Leveraging cross-domain experience for agentic problem solving

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024diversity (2024): Diversity empowers intelligence: Integrating expertise of software engineering agents

RQ3 / Representation Formats

Unstructured Text (125)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yu2026does (2026): Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chakraborty2025blaze (2025): BLAZE: Cross-language and cross-project bug localization via dynamic chunking and hard example learning

  • chang2025bridging (2025): Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models

  • chen2025locagent (2025): Locagent: Graph-guided llm agents for code localization

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • dai2025lita (2025): Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

  • ehrlich2025codemonkeys (2025): Codemonkeys: Scaling test-time compute for software engineering

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • gupta2025sacl (2025): SACL: Understanding and Combating Textual Bias in Code Retrieval with Semantic-Augmented Reranking and Localization

  • huang2025seeing (2025): Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • jiang2025cosil (2025): CoSIL: Software Issue Localization via LLM-Driven Code Repository Graph Searching

  • jiang2025putting (2025): Putting It All into Context: Simplifying Agents with LCLMs

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025infcode (2025): InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025sorft (2025): Sorft: Issue resolving with subtask-oriented reinforced fine-tuning

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • pabba2025semagent (2025): SemAgent: A Semantics Aware Program Repair Agent

  • pan2025codecor (2025): Codecor: An llm-based self-reflective multi-agent framework for code generation

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • reddy2025swerank (2025): SweRank: Software Issue Localization with Code Ranking

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • ruan2025specrover (2025): Specrover: Code intent extraction via llms

  • shi2025aime (2025): Aime: Towards Fully-Autonomous Multi-Agent Framework

  • shum2025swe (2025): SWE-RM: Execution-free Feedback For Software Engineering Agents

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • soni2025coding (2025): Coding Agents with Multimodal Browsing are Generalist Problem Solvers

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • su2025learn (2025): Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • tao2025code (2025): Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • tawosi2025meta (2025): Meta-RAG on Large Codebases Using Code Summarization

  • vaghasiya2025corethink (2025): CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs

  • vinh2025repeton (2025): Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wang2025improving (2025): Improving Code Localization with Repository Memory

  • wang2025mcts (2025): Mcts-refined cot: High-quality fine-tuning data for llm-based repository issue resolution

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025swe (2025): Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • xiong2025think (2025): Think-Search-Patch: A Retrieval-Augmented Reasoning Framework for Repository-Level Code Repair

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • yu2025utboost (2025): Utboost: Rigorous evaluation of coding agents on swe-bench

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025cast (2025): cast: Enhancing code retrieval-augmented generation with structural chunking via abstract syntax tree

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhang2025sealign (2025): Sealign: Alignment training for software engineering agent

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • arora2024masai (2024): Masai: Modular architecture for software-engineering ai agents

  • brown2024large (2024): Large language monkeys: Scaling inference compute with repeated sampling

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • cheshkov2024exploring (2024): Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • li2024codetree (2024): Codetree: Agent-guided tree search for code generation with large language models

  • lin2024llms (2024): Llms as continuous learners: Improving the reproduction of defective code in software issues

  • liu2024codexembed (2024): Codexembed: A generalist embedding model family for multiligual and multi-task code retrieval

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • ma2024repository (2024): Repository Structure-Aware Training Makes SLMs Better Issue Resolver

  • ouyang2024repograph (2024): Repograph: Enhancing ai software engineering with repository-level code graph

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • rombaut2024watson (2024): Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents

  • suresh2024cornstack (2024): CoRNStack: High-quality contrastive data for better code retrieval and reranking

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024autocoderover (2024): Autocoderover: Autonomous program improvement

  • zhang2024codev (2024): CodeV: Issue Resolving with Visual Data

  • zhang2024diversity (2024): Diversity empowers intelligence: Integrating expertise of software engineering agents

Structured Text (72)

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026empowering (2026): Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • yu2026does (2026): Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chakraborty2025blaze (2025): BLAZE: Cross-language and cross-project bug localization via dynamic chunking and hard example learning

  • chen2025locagent (2025): Locagent: Graph-guided llm agents for code localization

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • gupta2025sacl (2025): SACL: Understanding and Combating Textual Bias in Code Retrieval with Semantic-Augmented Reranking and Localization

  • jiang2025cosil (2025): CoSIL: Software Issue Localization via LLM-Driven Code Repository Graph Searching

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • liu2025graphlocator (2025): GraphLocator: Graph-guided Causal Reasoning for Issue Localization

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • pabba2025semagent (2025): SemAgent: A Semantics Aware Program Repair Agent

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • shum2025swe (2025): SWE-RM: Execution-free Feedback For Software Engineering Agents

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025agent (2025): Agent kb: Leveraging cross-domain experience for agentic problem solving

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tawosi2025meta (2025): Meta-RAG on Large Codebases Using Code Summarization

  • vinh2025repeton (2025): Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wang2025extracting (2025): Extracting Conceptual Knowledge to Locate Software Issues

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • xiong2025think (2025): Think-Search-Patch: A Retrieval-Augmented Reasoning Framework for Repository-Level Code Repair

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • lin2024llms (2024): Llms as continuous learners: Improving the reproduction of defective code in software issues

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

Graph (22)

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • chakraborty2025blaze (2025): BLAZE: Cross-language and cross-project bug localization via dynamic chunking and hard example learning

  • chang2025bridging (2025): Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models

  • chen2025locagent (2025): Locagent: Graph-guided llm agents for code localization

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • jiang2025cosil (2025): CoSIL: Software Issue Localization via LLM-Driven Code Repository Graph Searching

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • liu2025graphlocator (2025): GraphLocator: Graph-guided Causal Reasoning for Issue Localization

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • tao2025code (2025): Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • wang2025improving (2025): Improving Code Localization with Repository Memory

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • zhang2025sealign (2025): Sealign: Alignment training for software engineering agent

  • liu2024codexgraph (2024): Codexgraph: Bridging large language models and code repositories via code graph databases

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ouyang2024repograph (2024): Repograph: Enhancing ai software engineering with repository-level code graph

RQ4 / Knowledge Retrieval

Tool Invocation (106)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026empowering (2026): Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chen2025locagent (2025): Locagent: Graph-guided llm agents for code localization

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • dai2025lita (2025): Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

  • ehrlich2025codemonkeys (2025): Codemonkeys: Scaling test-time compute for software engineering

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • jiang2025cosil (2025): CoSIL: Software Issue Localization via LLM-Driven Code Repository Graph Searching

  • jiang2025putting (2025): Putting It All into Context: Simplifying Agents with LCLMs

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025infcode (2025): InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • liu2025graphlocator (2025): GraphLocator: Graph-guided Causal Reasoning for Issue Localization

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • pabba2025semagent (2025): SemAgent: A Semantics Aware Program Repair Agent

  • pan2025codecor (2025): Codecor: An llm-based self-reflective multi-agent framework for code generation

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • ruan2025specrover (2025): Specrover: Code intent extraction via llms

  • shi2025aime (2025): Aime: Towards Fully-Autonomous Multi-Agent Framework

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • soni2025coding (2025): Coding Agents with Multimodal Browsing are Generalist Problem Solvers

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • vaghasiya2025corethink (2025): CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs

  • vinh2025repeton (2025): Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wang2025improving (2025): Improving Code Localization with Repository Memory

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • yu2025utboost (2025): Utboost: Rigorous evaluation of coding agents on swe-bench

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • arora2024masai (2024): Masai: Modular architecture for software-engineering ai agents

  • brown2024large (2024): Large language monkeys: Scaling inference compute with repeated sampling

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • cheshkov2024exploring (2024): Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • lin2024llms (2024): Llms as continuous learners: Improving the reproduction of defective code in software issues

  • liu2024codexgraph (2024): Codexgraph: Bridging large language models and code repositories via code graph databases

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • ouyang2024repograph (2024): Repograph: Enhancing ai software engineering with repository-level code graph

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024autocoderover (2024): Autocoderover: Autonomous program improvement

Direct Prompt (95)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • xiang2026empowering (2026): Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yu2026does (2026): Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chakraborty2025blaze (2025): BLAZE: Cross-language and cross-project bug localization via dynamic chunking and hard example learning

  • chang2025bridging (2025): Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • dai2025lita (2025): Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

  • ehrlich2025codemonkeys (2025): Codemonkeys: Scaling test-time compute for software engineering

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • gupta2025sacl (2025): SACL: Understanding and Combating Textual Bias in Code Retrieval with Semantic-Augmented Reranking and Localization

  • huang2025seeing (2025): Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • jiang2025cosil (2025): CoSIL: Software Issue Localization via LLM-Driven Code Repository Graph Searching

  • jiang2025putting (2025): Putting It All into Context: Simplifying Agents with LCLMs

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • liu2025graphlocator (2025): GraphLocator: Graph-guided Causal Reasoning for Issue Localization

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025sorft (2025): Sorft: Issue resolving with subtask-oriented reinforced fine-tuning

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • pabba2025semagent (2025): SemAgent: A Semantics Aware Program Repair Agent

  • pan2025codecor (2025): Codecor: An llm-based self-reflective multi-agent framework for code generation

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • ruan2025specrover (2025): Specrover: Code intent extraction via llms

  • shum2025swe (2025): SWE-RM: Execution-free Feedback For Software Engineering Agents

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025agent (2025): Agent kb: Leveraging cross-domain experience for agentic problem solving

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • tao2025code (2025): Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • vinh2025repeton (2025): Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wang2025mcts (2025): Mcts-refined cot: High-quality fine-tuning data for llm-based repository issue resolution

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025swe (2025): Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • xiong2025think (2025): Think-Search-Patch: A Retrieval-Augmented Reasoning Framework for Repository-Level Code Repair

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • yu2025utboost (2025): Utboost: Rigorous evaluation of coding agents on swe-bench

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • cheshkov2024exploring (2024): Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • rombaut2024watson (2024): Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024autocoderover (2024): Autocoderover: Autonomous program improvement

  • zhang2024codev (2024): CodeV: Issue Resolving with Visual Data

  • zhang2024diversity (2024): Diversity empowers intelligence: Integrating expertise of software engineering agents

Parametric Injection (52)

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chakraborty2025blaze (2025): BLAZE: Cross-language and cross-project bug localization via dynamic chunking and hard example learning

  • chang2025bridging (2025): Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models

  • chen2025locagent (2025): Locagent: Graph-guided llm agents for code localization

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • ma2025sorft (2025): Sorft: Issue resolving with subtask-oriented reinforced fine-tuning

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • reddy2025swerank (2025): SweRank: Software Issue Localization with Code Ranking

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • shum2025swe (2025): SWE-RM: Execution-free Feedback For Software Engineering Agents

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • su2025learn (2025): Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tao2025code (2025): Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • wang2025mcts (2025): Mcts-refined cot: High-quality fine-tuning data for llm-based repository issue resolution

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025swe (2025): Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • xiong2025think (2025): Think-Search-Patch: A Retrieval-Augmented Reasoning Framework for Repository-Level Code Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhang2025sealign (2025): Sealign: Alignment training for software engineering agent

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • liu2024codexembed (2024): Codexembed: A generalist embedding model family for multiligual and multi-task code retrieval

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • ma2024repository (2024): Repository Structure-Aware Training Makes SLMs Better Issue Resolver

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • suresh2024cornstack (2024): CoRNStack: High-quality contrastive data for better code retrieval and reranking

Retrieval-Augmented Generation (47)

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • gupta2025sacl (2025): SACL: Understanding and Combating Textual Bias in Code Retrieval with Semantic-Augmented Reranking and Localization

  • huang2025seeing (2025): Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing

  • jiang2025cosil (2025): CoSIL: Software Issue Localization via LLM-Driven Code Repository Graph Searching

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • reddy2025swerank (2025): SweRank: Software Issue Localization with Code Ranking

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • shi2025aime (2025): Aime: Towards Fully-Autonomous Multi-Agent Framework

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • su2025learn (2025): Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments

  • tang2025agent (2025): Agent kb: Leveraging cross-domain experience for agentic problem solving

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • tao2025code (2025): Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • tawosi2025meta (2025): Meta-RAG on Large Codebases Using Code Summarization

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wang2025extracting (2025): Extracting Conceptual Knowledge to Locate Software Issues

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • xiong2025think (2025): Think-Search-Patch: A Retrieval-Augmented Reasoning Framework for Repository-Level Code Repair

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • zhang2025cast (2025): cast: Enhancing code retrieval-augmented generation with structural chunking via abstract syntax tree

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • brown2024large (2024): Large language monkeys: Scaling inference compute with repeated sampling

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • li2024codetree (2024): Codetree: Agent-guided tree search for code generation with large language models

  • lin2024llms (2024): Llms as continuous learners: Improving the reproduction of defective code in software issues

  • liu2024codexembed (2024): Codexembed: A generalist embedding model family for multiligual and multi-task code retrieval

  • liu2024codexgraph (2024): Codexgraph: Bridging large language models and code repositories via code graph databases

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ouyang2024repograph (2024): Repograph: Enhancing ai software engineering with repository-level code graph

  • suresh2024cornstack (2024): CoRNStack: High-quality contrastive data for better code retrieval and reranking

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

RQ5 / Usage Stages

Localization (117)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026empowering (2026): Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chakraborty2025blaze (2025): BLAZE: Cross-language and cross-project bug localization via dynamic chunking and hard example learning

  • chang2025bridging (2025): Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models

  • chen2025locagent (2025): Locagent: Graph-guided llm agents for code localization

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • dai2025lita (2025): Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

  • ehrlich2025codemonkeys (2025): Codemonkeys: Scaling test-time compute for software engineering

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • gupta2025sacl (2025): SACL: Understanding and Combating Textual Bias in Code Retrieval with Semantic-Augmented Reranking and Localization

  • huang2025seeing (2025): Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • jiang2025cosil (2025): CoSIL: Software Issue Localization via LLM-Driven Code Repository Graph Searching

  • jiang2025putting (2025): Putting It All into Context: Simplifying Agents with LCLMs

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025infcode (2025): InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • liu2025graphlocator (2025): GraphLocator: Graph-guided Causal Reasoning for Issue Localization

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025sorft (2025): Sorft: Issue resolving with subtask-oriented reinforced fine-tuning

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • pabba2025semagent (2025): SemAgent: A Semantics Aware Program Repair Agent

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • reddy2025swerank (2025): SweRank: Software Issue Localization with Code Ranking

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • ruan2025specrover (2025): Specrover: Code intent extraction via llms

  • shi2025aime (2025): Aime: Towards Fully-Autonomous Multi-Agent Framework

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • soni2025coding (2025): Coding Agents with Multimodal Browsing are Generalist Problem Solvers

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025agent (2025): Agent kb: Leveraging cross-domain experience for agentic problem solving

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • tao2025code (2025): Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • tawosi2025meta (2025): Meta-RAG on Large Codebases Using Code Summarization

  • vaghasiya2025corethink (2025): CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs

  • vinh2025repeton (2025): Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles

  • wang2025extracting (2025): Extracting Conceptual Knowledge to Locate Software Issues

  • wang2025improving (2025): Improving Code Localization with Repository Memory

  • wang2025mcts (2025): Mcts-refined cot: High-quality fine-tuning data for llm-based repository issue resolution

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wei2025swe (2025): Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • xiong2025think (2025): Think-Search-Patch: A Retrieval-Augmented Reasoning Framework for Repository-Level Code Repair

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025cast (2025): cast: Enhancing code retrieval-augmented generation with structural chunking via abstract syntax tree

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • arora2024masai (2024): Masai: Modular architecture for software-engineering ai agents

  • brown2024large (2024): Large language monkeys: Scaling inference compute with repeated sampling

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • li2024codetree (2024): Codetree: Agent-guided tree search for code generation with large language models

  • liu2024codexembed (2024): Codexembed: A generalist embedding model family for multiligual and multi-task code retrieval

  • liu2024codexgraph (2024): Codexgraph: Bridging large language models and code repositories via code graph databases

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • ouyang2024repograph (2024): Repograph: Enhancing ai software engineering with repository-level code graph

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • rombaut2024watson (2024): Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents

  • suresh2024cornstack (2024): CoRNStack: High-quality contrastive data for better code retrieval and reranking

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024autocoderover (2024): Autocoderover: Autonomous program improvement

  • zhang2024codev (2024): CodeV: Issue Resolving with Visual Data

Patch Generation (103)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yu2026does (2026): Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • dai2025lita (2025): Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

  • ehrlich2025codemonkeys (2025): Codemonkeys: Scaling test-time compute for software engineering

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • gupta2025sacl (2025): SACL: Understanding and Combating Textual Bias in Code Retrieval with Semantic-Augmented Reranking and Localization

  • huang2025seeing (2025): Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • jiang2025cosil (2025): CoSIL: Software Issue Localization via LLM-Driven Code Repository Graph Searching

  • jiang2025putting (2025): Putting It All into Context: Simplifying Agents with LCLMs

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025infcode (2025): InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025sorft (2025): Sorft: Issue resolving with subtask-oriented reinforced fine-tuning

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • pabba2025semagent (2025): SemAgent: A Semantics Aware Program Repair Agent

  • pan2025codecor (2025): Codecor: An llm-based self-reflective multi-agent framework for code generation

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • ruan2025specrover (2025): Specrover: Code intent extraction via llms

  • shi2025aime (2025): Aime: Towards Fully-Autonomous Multi-Agent Framework

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • soni2025coding (2025): Coding Agents with Multimodal Browsing are Generalist Problem Solvers

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • su2025learn (2025): Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025agent (2025): Agent kb: Leveraging cross-domain experience for agentic problem solving

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • tao2025code (2025): Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • tawosi2025meta (2025): Meta-RAG on Large Codebases Using Code Summarization

  • vaghasiya2025corethink (2025): CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs

  • wang2025mcts (2025): Mcts-refined cot: High-quality fine-tuning data for llm-based repository issue resolution

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025swe (2025): Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • xiong2025think (2025): Think-Search-Patch: A Retrieval-Augmented Reasoning Framework for Repository-Level Code Repair

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025cast (2025): cast: Enhancing code retrieval-augmented generation with structural chunking via abstract syntax tree

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • arora2024masai (2024): Masai: Modular architecture for software-engineering ai agents

  • brown2024large (2024): Large language monkeys: Scaling inference compute with repeated sampling

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • cheshkov2024exploring (2024): Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • li2024codetree (2024): Codetree: Agent-guided tree search for code generation with large language models

  • liu2024codexembed (2024): Codexembed: A generalist embedding model family for multiligual and multi-task code retrieval

  • liu2024codexgraph (2024): Codexgraph: Bridging large language models and code repositories via code graph databases

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • ouyang2024repograph (2024): Repograph: Enhancing ai software engineering with repository-level code graph

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024autocoderover (2024): Autocoderover: Autonomous program improvement

  • zhang2024codev (2024): CodeV: Issue Resolving with Visual Data

Patch Validation (89)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • yu2026does (2026): Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • dai2025lita (2025): Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

  • ehrlich2025codemonkeys (2025): Codemonkeys: Scaling test-time compute for software engineering

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • jiang2025putting (2025): Putting It All into Context: Simplifying Agents with LCLMs

  • lei2025infantagent (2025): InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • li2025infcode (2025): InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • pan2025codecor (2025): Codecor: An llm-based self-reflective multi-agent framework for code generation

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • ruan2025specrover (2025): Specrover: Code intent extraction via llms

  • shum2025swe (2025): SWE-RM: Execution-free Feedback For Software Engineering Agents

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • soni2025coding (2025): Coding Agents with Multimodal Browsing are Generalist Problem Solvers

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tang2025synfix (2025): SynFix: Dependency-aware program repair via RelationGraph analysis

  • vaghasiya2025corethink (2025): CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs

  • vinh2025repeton (2025): Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles

  • wang2025mcts (2025): Mcts-refined cot: High-quality fine-tuning data for llm-based repository issue resolution

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025demystifying (2025): Demystifying llm-based software engineering agents

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • xiao2025improving (2025): Improving the efficiency of LLM agent systems through trajectory reduction

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • yu2025utboost (2025): Utboost: Rigorous evaluation of coding agents on swe-bench

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • arora2024masai (2024): Masai: Modular architecture for software-engineering ai agents

  • brown2024large (2024): Large language monkeys: Scaling inference compute with repeated sampling

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • cheshkov2024exploring (2024): Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

  • gautam2024supercoder2 (2024): SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

  • zhang2024autocoderover (2024): Autocoderover: Autonomous program improvement

  • zhang2024diversity (2024): Diversity empowers intelligence: Integrating expertise of software engineering agents

Model Training (52)

  • kon2026swe (2026): SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • raghavendra2026agentic (2026): Agentic Rubrics as Contextual Verifiers for SWE Agents

  • song2026swe (2026): SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • tao2026swe (2026): Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chakraborty2025blaze (2025): BLAZE: Cross-language and cross-project bug localization via dynamic chunking and hard example learning

  • chang2025bridging (2025): Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models

  • chen2025locagent (2025): Locagent: Graph-guided llm agents for code localization

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • golubev2025training (2025): Training long-context, multi-turn software engineering agents with reinforcement learning

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • liu2025context (2025): Context as a tool: Context management for long-horizon swe-agents

  • ma2025sorft (2025): Sorft: Issue resolving with subtask-oriented reinforced fine-tuning

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • ma2025tool (2025): Tool-integrated reinforcement learning for repo deep search

  • rastogi2025devstral (2025): Devstral: Fine-tuning Language Models for Coding Agent Applications

  • reddy2025swerank (2025): SweRank: Software Issue Localization with Code Ranking

  • reddy2025swerank+ (2025): SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • shum2025swe (2025): SWE-RM: Execution-free Feedback For Software Engineering Agents

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • su2025learn (2025): Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025boosting (2025): Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • tang2025co (2025): Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • tao2025code (2025): Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • wang2025mcts (2025): Mcts-refined cot: High-quality fine-tuning data for llm-based repository issue resolution

  • wang2025practitioner (2025): A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025swe (2025): Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • xie2025swe (2025): Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • xiong2025think (2025): Think-Search-Patch: A Retrieval-Augmented Reasoning Framework for Repository-Level Code Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025building (2025): Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • zainullina2025guided (2025): Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • zeng2025satori (2025): Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • zhang2025one (2025): One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • zhang2025sealign (2025): Sealign: Alignment training for software engineering agent

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • liu2024codexembed (2024): Codexembed: A generalist embedding model family for multiligual and multi-task code retrieval

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • ma2024repository (2024): Repository Structure-Aware Training Makes SLMs Better Issue Resolver

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

  • suresh2024cornstack (2024): CoRNStack: High-quality contrastive data for better code retrieval and reranking

Reproduction Test Generation (44)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • liu2026dynamic (2026): Dynamic analysis enhances issue resolution

  • soni2026swe (2026): SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • xie2026arkeval (2026): ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chen2025old (2025): When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • chen2025prometheus (2025): Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • ehrlich2025codemonkeys (2025): Codemonkeys: Scaling test-time compute for software engineering

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • huang2025seeing (2025): Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • jiang2025putting (2025): Putting It All into Context: Simplifying Agents with LCLMs

  • li2025infcode (2025): InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ma2025thinking (2025): Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • mu2025experepair (2025): EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • nashid2025issue2test (2025): Issue2test: Generating reproducing test cases from issue reports

  • pabba2025semagent (2025): SemAgent: A Semantics Aware Program Repair Agent

  • ruan2025specrover (2025): Specrover: Code intent extraction via llms

  • sohrabizadeh2025nemotron (2025): Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • vinh2025repeton (2025): Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles

  • wang2025aegis (2025): Aegis: An agent-based framework for bug reproduction from issue descriptions

  • wang2025swe (2025): SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • wei2025swe (2025): Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution

  • wei2025toward (2025): Toward training superintelligent software agents through self-play swe-rl

  • yang2025enhancing (2025): Enhancing repository-level software repair via repository-aware knowledge graphs

  • yang2025kgcompass (2025): KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • yang2025kimi (2025): Kimi-dev: Agentless training as skill prior for swe-agents

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • yu2025orcaloca (2025): Orcaloca: An llm agent framework for software issue localization

  • yu2025utboost (2025): Utboost: Rigorous evaluation of coding agents on swe-bench

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • arora2024masai (2024): Masai: Modular architecture for software-engineering ai agents

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • lin2024llms (2024): Llms as continuous learners: Improving the reproduction of defective code in software issues

  • liu2024marscode (2024): Marscode agent: Ai-native automated bug fixing

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

Task Planning (38)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • ding2026swe (2026): SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • liu2026architecture (2026): Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • aggarwal2025dars (2025): Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • ahmed2025otter (2025): Otter: Generating tests from issues to validate swe patches

  • cao2025siadafix (2025): SIADAFIX: issue description response for adaptive program repair

  • chen2025swe (2025): Swe-exp: Experience-driven software issue resolution

  • da2025agent (2025): Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • dai2025lita (2025): Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

  • gandhi2025agents (2025): When agents go astray: Course-correcting swe agents with prms

  • gao2025trae (2025): Trae agent: An llm-based agent for software engineering with test-time scaling

  • li2025patchpilot (2025): PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • li2025swe (2025): Swe-debate: Competitive multi-agent debate for software issue resolution

  • lin2025se (2025): Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • lindenbauer2025complexity (2025): The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • ma2025alibaba (2025): Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • ouyang2025reasoningbank (2025): Reasoningbank: Scaling agent self-evolving with reasoning memory

  • pimpale2025forecasting (2025): Forecasting Frontier Language Model Agent Capabilities

  • sonwane2025bugpilot (2025): Bugpilot: Complex bug generation for efficient learning of swe skills

  • sun2025scaling (2025): Scaling long-horizon llm agent via context-folding

  • tang2025agent (2025): Agent kb: Leveraging cross-domain experience for agentic problem solving

  • wang2025seamlessflow (2025): SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • wong2025confucius (2025): Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • wu2025git (2025): Git context controller: Manage the context of llm-based agents like git

  • xia2025live (2025): Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • yang2025lingxi (2025): Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • zhang2025darwin (2025): Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • zhou2025tom (2025): Tom-swe: User mental modeling for software engineering agents

  • zhu2025training (2025): Training Versatile Coding Agents in Synthetic Environments

  • antoniades2024swe (2024): Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • chen2024coder (2024): Coder: Issue resolving with multi-agent and task graphs

  • lei2024infant (2024): Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • ma2024lingma (2024): Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • phan2024hyperagent (2024): Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • tao2024magis (2024): Magis: Llm-based multi-agent framework for github issue resolution

  • wang2024openhands (2024): Openhands: An open platform for ai software developers as generalist agents

  • yang2024swe (2024): Swe-agent: Agent-computer interfaces enable automated software engineering

Environment Building (10)

  • chen2026beyondswe (2026): BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • luo2026closing (2026): Closing the Loop: Universal Repository Representation with RPG-Encoder

  • xiang2026evaluating (2026): Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • yuan2026swe (2026): SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • copet2025cwm (2025): Cwm: An open-weights llm for research on code generation with world models

  • guo2025swe (2025): SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • jain2025r2e (2025): R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • yang2025swe (2025): Swe-smith: Scaling data for software engineering agents

  • zeng2025skywork (2025): Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • pan2024training (2024): Training software engineering agents and verifiers with swe-gym

Complete Registry

BeyondSWE: Can Current Code Agent Survive Beyond Single-Repo Bug Fixing?

  • Bib key: chen2026beyondswe
  • Year: 2026
  • Authors: Chen, Guoxin; Meng, Fanzhe; Zhao, Jiale; Li, Minghao; Cheng, Daixuan; Song, Huatong; Chen, Jie; Lin, Yuzhi; Chen, Hui; Zhao, Xin; others
  • Venue: arXiv preprint arXiv:2603.03194
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: discussion, findings

  • RQ1 / Knowledge Layers: Background Knowledge, Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Design Architecture Knowledge, Development Tool Knowledge, Domain-specific Knowledge, Existing Code Knowledge, Experiential Knowledge, Technology Stack Knowledge

  • RQ2 / Knowledge Sources: Codebase, Document, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Environment Building, Localization, Patch Generation, Patch Validation, Reproduction Test Generation, Task Planning

SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via Continuous Integration

  • Bib key: chen2026swe
  • Year: 2026
  • Authors: Chen, Jialong; Xu, Xander; Wei, Hu; Chen, Chuan; Zhao, Bing
  • Venue: arXiv preprint arXiv:2603.03823
  • Reference role: future_work_case
  • Screening status: included
  • Citation sections: discussion

EvoClaw: Evaluating AI Agents on Continuous Software Evolution

  • Bib key: deng2026evoclaw
  • Year: 2026
  • Authors: Deng, Gangda; Chen, Zhaoling; Yu, Zhongming; Fan, Haoyang; Liu, Yuhong; Yang, Yuxin; Parikh, Dhruv; Kannan, Rajgopal; Cong, Le; Wang, Mengdi; others
  • Venue: arXiv preprint arXiv:2603.13428
  • Reference role: future_work_case
  • Screening status: excluded
  • Citation sections: discussion

SWE-Replay: Efficient Test-Time Scaling for Software Engineering Agents

  • Bib key: ding2026swe
  • Year: 2026
  • Authors: Ding, Yifeng; Zhang, Lingming
  • Venue: arXiv preprint arXiv:2601.22129
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?

  • Bib key: han2026swe
  • Year: 2026
  • Authors: Han, Tingxu; Zhang, Yi; Song, Wei; Fang, Chunrong; Chen, Zhenyu; Sun, Youcheng; Hu, Lijie
  • Venue: arXiv preprint arXiv:2603.15401
  • Reference role: future_work_case
  • Screening status: excluded
  • Citation sections: discussion

SWE-Prot$\backslash$'eg$\backslash$'e: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

  • Bib key: kon2026swe
  • Year: 2026
  • Authors: Kon, Patrick Tser Jern; Pradeep, Archana; Chen, Ang; Ellis, Alexander P; Hunt, Warren; Wang, Zijian; Yang, John; Thompson, Samuel
  • Venue: arXiv preprint arXiv:2602.22124
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Document, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation

Advances and Frontiers of LLM-based Issue Resolution in Software Engineering: A Comprehensive Survey

  • Bib key: li2026advancesFrontiers
  • Year: 2026
  • Authors: Li, Caihua; Guo, Lianghong; Wang, Yanlin; Guo, Daya; Tao, Wei; Shan, Zhenyu; Liu, Mingwei; Chen, Jiachi; Song, Haoyu; Tang, Duyu; others
  • Venue: arXiv preprint arXiv:2601.11655
  • Reference role: related_survey
  • Screening status: excluded
  • Citation sections: background, introduction

Architecture-Aware Multi-Design Generation for Repository-Level Feature Addition

  • Bib key: liu2026architecture
  • Year: 2026
  • Authors: Liu, Mingwei; Chen, Zhenxi; Pei, Zheng; Wang, Zihao; Wang, Yanlin; Zheng, Zibin
  • Venue: arXiv preprint arXiv:2603.01814
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Design Architecture Knowledge, Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

Dynamic analysis enhances issue resolution

  • Bib key: liu2026dynamic
  • Year: 2026
  • Authors: Liu, Mingwei; Wang, Zihao; Chen, Zhenxi; Pei, Zheng; Wang, Yanlin; Zheng, Zibin
  • Venue: arXiv preprint arXiv:2603.22048
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings, introduction

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation

Closing the Loop: Universal Repository Representation with RPG-Encoder

  • Bib key: luo2026closing
  • Year: 2026
  • Authors: Luo, Jane; Yin, Chengyu; Zhang, Xin; Li, Qingtao; Liu, Steven; Huang, Yiming; Wu, Jie; Liu, Hao; Huang, Yangyu; Kang, Yu; others
  • Venue: arXiv preprint arXiv:2602.02084
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Design Architecture Knowledge, Development Tool Knowledge, Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Expert, Repository History

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Environment Building, Localization, Patch Generation, Task Planning

Agentic Rubrics as Contextual Verifiers for SWE Agents

  • Bib key: raghavendra2026agentic
  • Year: 2026
  • Authors: Raghavendra, Mohit; Gunjal, Anisha; Liu, Bing; He, Yunzhong
  • Venue: arXiv preprint arXiv:2601.04171
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Design Architecture Knowledge, Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Document, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Model Training, Patch Validation

SWE-Master: Unleashing the Potential of Software Engineering Agents via Post-Training

  • Bib key: song2026swe
  • Year: 2026
  • Authors: Song, Huatong; Huang, Lisheng; Sun, Shuang; Jiang, Jinhao; Le, Ran; Cheng, Daixuan; Chen, Guoxin; Hu, Yiwen; Chen, Zongchao; Jia, Yiming; others
  • Venue: arXiv preprint arXiv:2602.03411
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation

SWE-Tester: Training Open-Source LLMs for Issue Reproduction in Real-World Repositories

  • Bib key: soni2026swe
  • Year: 2026
  • Authors: Soni, Aditya Bharat; Ghosh, Rajat; Bhargava, Vaishnavi; Chen, Valerie; Dutta, Debojyoti
  • Venue: arXiv preprint arXiv:2601.13713
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Validation, Reproduction Test Generation

Swe-lego: Pushing the limits of supervised fine-tuning for software issue resolving

  • Bib key: tao2026swe
  • Year: 2026
  • Authors: Tao, Chaofan; Chen, Jierun; Jiang, Yuxin; Kou, Kaiqi; Wang, Shaowei; Wang, Ruoyu; Li, Xiaohui; Yang, Sidi; Du, Yiming; Dai, Jianbo; others
  • Venue: arXiv preprint arXiv:2601.01426
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert, Repository History

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation

MemGovern: Enhancing Code Agents through Learning from Governed Human Experiences

  • Bib key: wang2026memgovern
  • Year: 2026
  • Authors: Wang, Qihao; Cheng, Ziming; Zhang, Shuo; Liu, Fan; Xu, Rui; Lian, Heng; Wang, Kunyi; Yu, Xiaoming; Yin, Jianghao; Hu, Sen; others
  • Venue: arXiv preprint arXiv:2601.06789
  • Reference role: context_reference
  • Screening status: excluded
  • Citation sections: none

Empowering Autonomous Debugging Agents with Efficient Dynamic Analysis

  • Bib key: xiang2026empowering
  • Year: 2026
  • Authors: Xiang, Jiahong; Xu, Xiaoyang; Chu, Xiaopan; Tian, Hongliang; Zhang, Yuqun
  • Venue: arXiv preprint arXiv:2604.24212
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization

Evaluating and Improving Automated Repository-Level Rust Issue Resolution with LLM-based Agents

  • Bib key: xiang2026evaluating
  • Year: 2026
  • Authors: Xiang, Jiahong; He, Wenxiao; Wang, Xihua; Tian, Hongliang; Zhang, Yuqun
  • Venue: arXiv preprint arXiv:2602.22764
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: discussion, findings

  • RQ1 / Knowledge Layers: Background Knowledge, Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Programming Language Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Environment Building, Localization, Patch Generation, Patch Validation, Reproduction Test Generation

ArkEval: Benchmarking and Evaluating Automated CodeRepair for ArkTS

  • Bib key: xie2026arkeval
  • Year: 2026
  • Authors: Xie, Bang; Zhang, Senjian; Peng, Zhiyuan; Chen, Wei; Ying, Chenhao; Luo, Yuan
  • Venue: arXiv preprint arXiv:2602.08866
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: discussion, findings

  • RQ1 / Knowledge Layers: Background Knowledge, Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Programming Language Knowledge, Repository Evolution Knowledge, Technology Stack Knowledge

  • RQ2 / Knowledge Sources: Codebase, Document, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation

Does Pass Rate Tell the Whole Story? Evaluating Design Constraint Compliance in LLM-based Issue Resolution

  • Bib key: yu2026does
  • Year: 2026
  • Authors: Yu, Kai; Zhou, Zhenhao; Zeng, Junhao; Wang, Ying; Du, Xueying; Yuan, Zhiqiang; Liu, Junwei; Zhou, Ziyu; Wang, Yujia; Wang, Chong; others
  • Venue: arXiv preprint arXiv:2604.05955
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Design Architecture Knowledge, Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt

  • RQ5 / Usage Stages: Patch Generation, Patch Validation

SWE-MiniSandbox: Container-Free Reinforcement Learning for Building Software Engineering Agents

  • Bib key: yuan2026swe
  • Year: 2026
  • Authors: Yuan, Danlong; Wu, Wei; Wang, Zhengren; Zhao, Xueliang; Zhang, Huishuai; Zhao, Dongyan
  • Venue: arXiv preprint arXiv:2602.11210
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Environment Building, Localization, Model Training, Patch Generation, Patch Validation

Dars: Dynamic action re-sampling to enhance coding agent performance by adaptive tree traversal

  • Bib key: aggarwal2025dars
  • Year: 2025
  • Authors: Aggarwal, Vaibhav; Kamal, Ojasv; Japesh, Abhinav; Jin, Zhijing; Sch{\"o}lkopf, Bernhard
  • Venue: arXiv preprint arXiv:2503.14269
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation, Reproduction Test Generation, Task Planning

Otter: Generating tests from issues to validate swe patches

  • Bib key: ahmed2025otter
  • Year: 2025
  • Authors: Ahmed, Toufique; Ganhotra, Jatin; Pan, Rangeet; Shinnar, Avraham; Sinha, Saurabh; Hirzel, Martin
  • Venue: arXiv preprint arXiv:2502.05368
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert, Repository History

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Validation, Reproduction Test Generation, Task Planning

SIADAFIX: issue description response for adaptive program repair

  • Bib key: cao2025siadafix
  • Year: 2025
  • Authors: Cao, Xin; Yu, Nan
  • Venue: arXiv preprint arXiv:2510.16059
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Design Architecture Knowledge, Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert, Repository History

  • RQ2 / Extraction Methods: Manual Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation, Reproduction Test Generation, Task Planning

BLAZE: Cross-language and cross-project bug localization via dynamic chunking and hard example learning

  • Bib key: chakraborty2025blaze
  • Year: 2025
  • Authors: Chakraborty, Partha; Alfadel, Mahmoud; Nagappan, Meiyappan
  • Venue: IEEE Transactions on Software Engineering
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Repository History

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection

  • RQ5 / Usage Stages: Localization, Model Training

Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models

  • Bib key: chang2025bridging
  • Year: 2025
  • Authors: Chang, Jianming; Zhou, Xin; Wang, Lulu; Lo, David; Li, Bixin
  • Venue: arXiv preprint arXiv:2502.15292
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection

  • RQ5 / Usage Stages: Localization, Model Training

Locagent: Graph-guided llm agents for code localization

  • Bib key: chen2025locagent
  • Year: 2025
  • Authors: Chen, Zhaoling; Tang, Xiangru; Deng, Gangda; Wu, Fang; Wu, Jialong; Jiang, Zhiwei; Prasanna, Viktor; Cohan, Arman; Wang, Xingyao
  • Venue: arXiv preprint arXiv:2503.09089
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training

When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution

  • Bib key: chen2025old
  • Year: 2025
  • Authors: Chen, Yang; Ahmed, Toufique; Jabbarvand, Reyhaneh; Hirzel, Martin
  • Venue: arXiv preprint arXiv:2510.18270
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Validation, Reproduction Test Generation

Prometheus: Unified knowledge graphs for issue resolution in multilingual codebases

  • Bib key: chen2025prometheus
  • Year: 2025
  • Authors: Chen, Zimin; Pan, Yue; Lu, Siyu; Xu, Jiayi; Goues, Claire Le; Monperrus, Martin; Ye, He
  • Venue: arXiv preprint arXiv:2507.19942
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings, introduction

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation

Swe-exp: Experience-driven software issue resolution

  • Bib key: chen2025swe
  • Year: 2025
  • Authors: Chen, Silin; Lin, Shaoxin; Gu, Xiaodong; Shi, Yuling; Lian, Heng; Yun, Longfei; Chen, Dong; Sun, Weiguo; Cao, Lin; Wang, Qianxiang
  • Venue: arXiv preprint arXiv:2507.23361
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings, introduction

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Task Planning

Cwm: An open-weights llm for research on code generation with world models

  • Bib key: copet2025cwm
  • Year: 2025
  • Authors: Copet, Jade; Carbonneaux, Quentin; Cohen, Gal; Gehring, Jonas; Kahn, Jacob; Kossen, Jannik; Kreuk, Felix; McMilin, Emily; Meyer, Michel; Wei, Yuxiang; others
  • Venue: arXiv preprint arXiv:2510.02387
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Environment Building, Localization, Model Training, Patch Generation, Patch Validation

Agent-RLVR: Training Software Engineering Agents via Guidance and Environment Rewards

  • Bib key: da2025agent
  • Year: 2025
  • Authors: Da, Jeff; Wang, Clinton; Deng, Xiang; Ma, Yuntao; Barhate, Nikhil; Hendryx, Sean
  • Venue: arXiv preprint arXiv:2506.11425
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation, Task Planning

Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs

  • Bib key: dai2025lita
  • Year: 2025
  • Authors: Dai, Hankun; Wang, Maoquan; Qi, Mengnan; Zhang, Yikai; Jin, Zijian; Yao, Yongqiang; Huang, Yufan; Fu, Shengyu; Nallipogu, Elsie
  • Venue: arXiv preprint arXiv:2509.25873
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

Swe-bench pro: Can ai agents solve long-horizon software engineering tasks?

  • Bib key: deng2025swe
  • Year: 2025
  • Authors: Deng, Xiang; Da, Jeff; Pan, Edwin; He, Yannis Yiming; Ide, Charles; Garg, Kanak; Lauffer, Niklas; Park, Andrew; Pasari, Nitin; Rane, Chetan; others
  • Venue: arXiv preprint arXiv:2509.16941
  • Reference role: future_work_case
  • Screening status: excluded
  • Citation sections: discussion

Codemonkeys: Scaling test-time compute for software engineering

  • Bib key: ehrlich2025codemonkeys
  • Year: 2025
  • Authors: Ehrlich, Ryan; Brown, Bradley; Juravsky, Jordan; Clark, Ronald; R{\'e}, Christopher; Mirhoseini, Azalia
  • Venue: arXiv preprint arXiv:2501.14723
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation

CoRet: Improved Retriever for Code Editing

  • Bib key: fehr2025coret
  • Year: 2025
  • Authors: Fehr, Fabio; Sivaprasad, Prabhu Teja; Franceschi, Luca; Zappella, Giovanni
  • Venue: arXiv preprint arXiv:2505.24715
  • Reference role: context_reference
  • Screening status: excluded
  • Citation sections: none

When agents go astray: Course-correcting swe agents with prms

  • Bib key: gandhi2025agents
  • Year: 2025
  • Authors: Gandhi, Shubham; Tsay, Jason; Ganhotra, Jatin; Kate, Kiran; Rizk, Yara
  • Venue: arXiv preprint arXiv:2509.02360
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

Trae agent: An llm-based agent for software engineering with test-time scaling

  • Bib key: gao2025trae
  • Year: 2025
  • Authors: Gao, Pengfei; Tian, Zhao; Meng, Xiangxin; Wang, Xinchen; Hu, Ruida; Xiao, Yuanan; Liu, Yizhou; Zhang, Zhao; Chen, Junjie; Gao, Cuiyun; others
  • Venue: arXiv preprint arXiv:2507.23370
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation, Task Planning

Training long-context, multi-turn software engineering agents with reinforcement learning

  • Bib key: golubev2025training
  • Year: 2025
  • Authors: Golubev, Alexander; Trofimova, Maria; Polezhaev, Sergei; Badertdinov, Ibragim; Nekrashevich, Maksim; Shevtsov, Anton; Karasik, Simon; Abramov, Sergey; Andriushchenko, Andrei; Fisin, Filipp; others
  • Venue: arXiv preprint arXiv:2508.03501
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation

SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks

  • Bib key: guo2025swe
  • Year: 2025
  • Authors: Guo, Lianghong; Wang, Yanlin; Li, Caihua; Yang, Pengyu; Chen, Jiachi; Tao, Wei; Zou, Yingtian; Tang, Duyu; Zheng, Zibin
  • Venue: arXiv preprint arXiv:2506.10954
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Document, Dynamic Execution Results, Expert, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Environment Building, Patch Validation

SACL: Understanding and Combating Textual Bias in Code Retrieval with Semantic-Augmented Reranking and Localization

  • Bib key: gupta2025sacl
  • Year: 2025
  • Authors: Gupta, Dhruv; Lakshmy, Gayathri Ganesh; Xie, Yiqing
  • Venue: arXiv preprint arXiv:2506.20081
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation

  • RQ5 / Usage Stages: Localization, Patch Generation

Seeing is Fixing: Cross-Modal Reasoning with Multimodal LLMs for Visual Software Issue Fixing

  • Bib key: huang2025seeing
  • Year: 2025
  • Authors: Huang, Kai; Zhang, Jian; Xie, Xiaofei; Chen, Chunyang
  • Venue: arXiv preprint arXiv:2506.16136
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Design Architecture Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Document

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation

  • RQ5 / Usage Stages: Localization, Patch Generation, Reproduction Test Generation

R2e-gym: Procedural environments and hybrid verifiers for scaling open-weights swe agents

  • Bib key: jain2025r2e
  • Year: 2025
  • Authors: Jain, Naman; Singh, Jaskirat; Shetty, Manish; Zheng, Liang; Sen, Koushik; Stoica, Ion
  • Venue: arXiv preprint arXiv:2504.07164
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Background Knowledge, Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge, Technology Stack Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert, Repository History

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Environment Building, Localization, Model Training, Patch Generation, Patch Validation, Reproduction Test Generation

Agentic Software Issue Resolution with Large Language Models: A Survey

  • Bib key: jiang2025agenticSurvey
  • Year: 2025
  • Authors: Jiang, Zhonghao; Lo, David; Liu, Zhongxin
  • Venue: arXiv preprint arXiv:2512.22256
  • Reference role: related_survey
  • Screening status: excluded
  • Citation sections: background, introduction

CoSIL: Software Issue Localization via LLM-Driven Code Repository Graph Searching

  • Bib key: jiang2025cosil
  • Year: 2025
  • Authors: Jiang, Zhonghao; Ren, Xiaoxue; Yan, Meng; Jiang, Wei; Li, Yong; Liu, Zhongxin
  • Venue: arXiv preprint arXiv:2503.22424
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: background, findings, introduction

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation

Putting It All into Context: Simplifying Agents with LCLMs

  • Bib key: jiang2025putting
  • Year: 2025
  • Authors: Jiang, Mingjian; Ruan, Yangjun; Lastras, Luis; Kapanipathi, Pavan; Hashimoto, Tatsunori
  • Venue: arXiv preprint arXiv:2505.08120
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation

InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction

  • Bib key: lei2025infantagent
  • Year: 2025
  • Authors: Lei, Bin; Kang, Weitai; Zhang, Zijian; Chen, Winson; Xie, Xi; Zuo, Shan; Xie, Mimi; Payani, Ali; Hong, Mingyi; Yan, Yan; others
  • Venue: arXiv preprint arXiv:2505.10887
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation

Fea-bench: A benchmark for evaluating repository-level code generation for feature implementation

  • Bib key: li2025fea
  • Year: 2025
  • Authors: Li, Wei; Zhang, Xin; Guo, Zhongxin; Mao, Shaoguang; Luo, Wen; Peng, Guangyue; Huang, Yangyu; Wang, Houfeng; Li, Scarlett
  • Venue: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
  • Reference role: benchmark_background
  • Screening status: excluded
  • Citation sections: background, discussion

InfCode: Adversarial Iterative Refinement of Tests and Patches for Reliable Software Issue Resolution

  • Bib key: li2025infcode
  • Year: 2025
  • Authors: Li, KeFan; Wang, Mengfei; Zhang, Hengzhi; Li, Zhichao; Yuan, Yuan; Li, Mu; Gao, Xiang; Sun, Hailong; Hu, Chunming; Lv, Weifeng
  • Venue: arXiv preprint arXiv:2511.16004
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation

PatchPilot: A Cost-Efficient Software Engineering Agent with Early Attempts on Formal Verification

  • Bib key: li2025patchpilot
  • Year: 2025
  • Authors: Li, Hongwei; Tang, Yuheng; Wang, Shiqi; Guo, Wenbo
  • Venue: arXiv preprint arXiv:2502.02747
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert, Repository History

  • RQ2 / Extraction Methods: Manual Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation, Task Planning

Swe-debate: Competitive multi-agent debate for software issue resolution

  • Bib key: li2025swe
  • Year: 2025
  • Authors: Li, Han; Shi, Yuling; Lin, Shaoxin; Gu, Xiaodong; Lian, Heng; Wang, Xin; Jia, Yantao; Huang, Tao; Wang, Qianxiang
  • Venue: arXiv preprint arXiv:2507.23348
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Design Architecture Knowledge, Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

Se-agent: Self-evolution trajectory optimization in multi-step reasoning with llm-based agents

  • Bib key: lin2025se
  • Year: 2025
  • Authors: Lin, Jiaye; Guo, Yifu; Han, Yuzhen; Hu, Sen; Ni, Ziyi; Wang, Licheng; Chen, Mingguang; Liu, Hongzhang; Chen, Ronghao; He, Yangfan; others
  • Venue: arXiv preprint arXiv:2508.02085
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

  • Bib key: lindenbauer2025complexity
  • Year: 2025
  • Authors: Lindenbauer, Tobias; Slinko, Igor; Felder, Ludwig; Bogomolov, Egor; Zharov, Yaroslav
  • Venue: arXiv preprint arXiv:2508.21433
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

Context as a tool: Context management for long-horizon swe-agents

  • Bib key: liu2025context
  • Year: 2025
  • Authors: Liu, Shukai; Yang, Jian; Jiang, Bo; Li, Yizhi; Guo, Jinyang; Liu, Xianglong; Dai, Bryan
  • Venue: arXiv preprint arXiv:2512.22087
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation

An Empirical Study on Failures in Automated Issue Solving

  • Bib key: liu2025empirical
  • Year: 2025
  • Authors: Liu, Simiao; Liu, Fang; Li, Liehao; Tan, Xin; Zhu, Yinghao; Lian, Xiaoli; Zhang, Li
  • Venue: arXiv preprint arXiv:2509.13941
  • Reference role: context_reference
  • Screening status: excluded
  • Citation sections: none

GraphLocator: Graph-guided Causal Reasoning for Issue Localization

  • Bib key: liu2025graphlocator
  • Year: 2025
  • Authors: Liu, Wei; Peng, Chao; Gao, Pengfei; Liu, Aofan; Zhang, Wei; Zhao, Haiyan; Jin, Zhi
  • Venue: arXiv preprint arXiv:2512.22469
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Structured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization

Alibaba lingmaagent: Improving automated issue resolution via comprehensive repository exploration

  • Bib key: ma2025alibaba
  • Year: 2025
  • Authors: Ma, Yingwei; Yang, Qingping; Cao, Rongyu; Li, Binhua; Huang, Fei; Li, Yongbin
  • Venue: Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation, Task Planning

Sorft: Issue resolving with subtask-oriented reinforced fine-tuning

  • Bib key: ma2025sorft
  • Year: 2025
  • Authors: Ma, Zexiong; Peng, Chao; Gao, Pengfei; Meng, Xiangxin; Zou, Yanzhen; Xie, Bing
  • Venue: arXiv preprint arXiv:2502.20127
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Repository History

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation

Thinking longer, not larger: Enhancing software engineering agents via scaling test-time compute

  • Bib key: ma2025thinking
  • Year: 2025
  • Authors: Ma, Yingwei; Li, Yongbin; Dong, Yihong; Jiang, Xue; Li, Yanhao; Liu, Yue; Cao, Rongyu; Chen, Jue; Huang, Fei; Li, Binhua
  • Venue: 2025 40th IEEE/ACM International Conference on Automated Software Engineering (ASE)
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Validation, Reproduction Test Generation

  • Bib key: ma2025tool
  • Year: 2025
  • Authors: Ma, Zexiong; Peng, Chao; Zeng, Qunhong; Gao, Pengfei; Zou, Yanzhen; Xie, Bing
  • Venue: arXiv preprint arXiv:2508.03012
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert, Repository History

  • RQ2 / Extraction Methods: Manual Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training

EXPEREPAIR: Dual-Memory Enhanced LLM-based Repository-Level Program Repair

  • Bib key: mu2025experepair
  • Year: 2025
  • Authors: Mu, Fangwen; Wang, Junjie; Shi, Lin; Wang, Song; Li, Shoubin; Wang, Qing
  • Venue: arXiv preprint arXiv:2506.10484
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation

Issue2test: Generating reproducing test cases from issue reports

  • Bib key: nashid2025issue2test
  • Year: 2025
  • Authors: Nashid, Noor; Bouzenia, Islem; Pradel, Michael; Mesbah, Ali
  • Venue: arXiv preprint arXiv:2503.16320
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Design Architecture Knowledge, Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Reproduction Test Generation

Reasoningbank: Scaling agent self-evolving with reasoning memory

  • Bib key: ouyang2025reasoningbank
  • Year: 2025
  • Authors: Ouyang, Siru; Yan, Jun; Hsu, I; Chen, Yanfei; Jiang, Ke; Wang, Zifeng; Han, Rujun; Le, Long T; Daruki, Samira; Tang, Xiangru; others
  • Venue: arXiv preprint arXiv:2509.25140
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Task Planning

SemAgent: A Semantics Aware Program Repair Agent

  • Bib key: pabba2025semagent
  • Year: 2025
  • Authors: Pabba, Anvith; Mathai, Alex; Chakraborty, Anindya; Ray, Baishakhi
  • Venue: arXiv preprint arXiv:2506.16650
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Reproduction Test Generation

Codecor: An llm-based self-reflective multi-agent framework for code generation

  • Bib key: pan2025codecor
  • Year: 2025
  • Authors: Pan, Ruwei; Zhang, Hongyu; Liu, Chao
  • Venue: arXiv preprint arXiv:2501.07811
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge

  • RQ2 / Knowledge Sources: Dynamic Execution Results

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Patch Generation, Patch Validation

Forecasting Frontier Language Model Agent Capabilities

  • Bib key: pimpale2025forecasting
  • Year: 2025
  • Authors: Pimpale, Govind; H{\o}jmark, Axel; Scheurer, J{\'e}r{\'e}my; Hobbhahn, Marius
  • Venue: arXiv preprint arXiv:2502.15850
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

Swe-polybench: A multi-language benchmark for repository level evaluation of coding agents

  • Bib key: rashid2025swe
  • Year: 2025
  • Authors: Rashid, Muhammad Shihab; Bock, Christian; Zhuang, Yuan; Buchholz, Alexander; Esler, Tim; Valentin, Simon; Franceschi, Luca; Wistuba, Martin; Sivaprasad, Prabhu Teja; Kim, Woo Jung; others
  • Venue: arXiv preprint arXiv:2504.08703
  • Reference role: future_work_case
  • Screening status: excluded
  • Citation sections: discussion

Devstral: Fine-tuning Language Models for Coding Agent Applications

  • Bib key: rastogi2025devstral
  • Year: 2025
  • Authors: Rastogi, Abhinav; Yang, Adam; Jiang, Albert Q; Liu, Alexander H; Sablayrolles, Alexandre; H{\'e}liou, Am{\'e}lie; Martin, Am{\'e}lie; Agarwal, Anmol; Ehrenberg, Andy; Lo, Andy; others
  • Venue: arXiv preprint arXiv:2509.25193
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation

SweRank: Software Issue Localization with Code Ranking

  • Bib key: reddy2025swerank
  • Year: 2025
  • Authors: Reddy, Revanth Gangi; Suresh, Tarun; Doo, JaeHyeok; Liu, Ye; Nguyen, Xuan Phi; Zhou, Yingbo; Yavuz, Semih; Xiong, Caiming; Ji, Heng; Joty, Shafiq
  • Venue: arXiv preprint arXiv:2505.07849
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: background, findings, introduction

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Retrieval-Augmented Generation

  • RQ5 / Usage Stages: Localization, Model Training

SweRank+: Multilingual, Multi-Turn Code Ranking for Software Issue Localization

  • Bib key: reddy2025swerank+
  • Year: 2025
  • Authors: Reddy, Revanth Gangi; Liu, Ye; Zhao, Wenting; Doo, JaeHyeok; Suresh, Tarun; Lee, Daniel; Xiong, Caiming; Zhou, Yingbo; Yavuz, Semih; Joty, Shafiq
  • Venue: arXiv preprint arXiv:2512.20482
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Background Knowledge, Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge, Experiential Knowledge, Programming Language Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Expert, Repository History

  • RQ2 / Extraction Methods: Manual Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training

Specrover: Code intent extraction via llms

  • Bib key: ruan2025specrover
  • Year: 2025
  • Authors: Ruan, Haifeng; Zhang, Yuntong; Roychoudhury, Abhik
  • Venue: 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE)
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation

Aime: Towards Fully-Autonomous Multi-Agent Framework

  • Bib key: shi2025aime
  • Year: 2025
  • Authors: Shi, Yexuan; Wang, Mingyu; Cao, Yunxiang; Lai, Hongjie; Lan, Junjian; Han, Xin; Wang, Yu; Geng, Jie; Li, Zhenan; Xia, Zihao; others
  • Venue: arXiv preprint arXiv:2507.11988
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation

SWE-RM: Execution-free Feedback For Software Engineering Agents

  • Bib key: shum2025swe
  • Year: 2025
  • Authors: Shum, KaShun; Hui, Binyuan; Chen, Jiawei; Zhang, Lei; Yang, Jiaxi; Huang, Yuzhen; Lin, Junyang; He, Junxian; others
  • Venue: arXiv preprint arXiv:2512.21919
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection

  • RQ5 / Usage Stages: Model Training, Patch Validation

Nemotron-cortexa: Enhancing llm agents for software engineering tasks via improved localization and solution diversity

  • Bib key: sohrabizadeh2025nemotron
  • Year: 2025
  • Authors: Sohrabizadeh, Atefeh; Song, Jialin; Liu, Mingjie; Roy, Rajarshi; Lee, Chankyu; Raiman, Jonathan; Catanzaro, Bryan
  • Venue: Forty-second International Conference on Machine Learning
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation, Reproduction Test Generation

Coding Agents with Multimodal Browsing are Generalist Problem Solvers

  • Bib key: soni2025coding
  • Year: 2025
  • Authors: Soni, Aditya Bharat; Li, Boxuan; Wang, Xingyao; Chen, Valerie; Neubig, Graham
  • Venue: arXiv preprint arXiv:2506.03011
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation

Bugpilot: Complex bug generation for efficient learning of swe skills

  • Bib key: sonwane2025bugpilot
  • Year: 2025
  • Authors: Sonwane, Atharv; White, Isadora; Lee, Hyunji; Pereira, Matheus; Caccia, Lucas; Kim, Minseon; Shi, Zhengyan; Singh, Chinmay; Sordoni, Alessandro; C{\^o}t{\'e}, Marc-Alexandre; others
  • Venue: arXiv preprint arXiv:2510.19898
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation, Task Planning

Learn-by-interact: A data-centric framework for self-adaptive agents in realistic environments

  • Bib key: su2025learn
  • Year: 2025
  • Authors: Su, Hongjin; Sun, Ruoxi; Yoon, Jinsung; Yin, Pengcheng; Yu, Tao; Ar{\i}k, Sercan {\"O}
  • Venue: arXiv preprint arXiv:2501.10893
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge

  • RQ1 / Knowledge Types: Experiential Knowledge

  • RQ2 / Knowledge Sources: Document, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Retrieval-Augmented Generation

  • RQ5 / Usage Stages: Model Training, Patch Generation

Scaling long-horizon llm agent via context-folding

  • Bib key: sun2025scaling
  • Year: 2025
  • Authors: Sun, Weiwei; Lu, Miao; Ling, Zhan; Liu, Kang; Yao, Xuesong; Yang, Yiming; Chen, Jiecao
  • Venue: arXiv preprint arXiv:2510.11967
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Document, Dynamic Execution Results

  • RQ2 / Extraction Methods: Manual Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation, Reproduction Test Generation, Task Planning

Agent kb: Leveraging cross-domain experience for agentic problem solving

  • Bib key: tang2025agent
  • Year: 2025
  • Authors: Tang, Xiangru; Qin, Tianrui; Peng, Tianhao; Zhou, Ziyang; Shao, Daniel; Du, Tingting; Wei, Xinming; Xia, Peng; Wu, Fang; Zhu, He; others
  • Venue: arXiv preprint arXiv:2507.06229
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings, introduction

  • RQ1 / Knowledge Layers: Procedural Knowledge

  • RQ1 / Knowledge Types: Experiential Knowledge

  • RQ2 / Knowledge Sources: Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction

  • RQ3 / Representation Formats: Structured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation

  • RQ5 / Usage Stages: Localization, Patch Generation, Task Planning

Boosting Open-Source LLMs for Program Repair via Reasoning Transfer and LLM-Guided Reinforcement Learning

  • Bib key: tang2025boosting
  • Year: 2025
  • Authors: Tang, Xunzhu; Klein, Jacques; Bissyand{\'e}, Tegawend{\'e} F
  • Venue: arXiv preprint arXiv:2506.03921
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation

Co-PatcheR: Collaborative Software Patching with Component (s)-specific Small Reasoning Models

  • Bib key: tang2025co
  • Year: 2025
  • Authors: Tang, Yuheng; Li, Hongwei; Zhu, Kaijie; Yang, Michael; Ding, Yangruibo; Guo, Wenbo
  • Venue: arXiv preprint arXiv:2505.18955
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation

SynFix: Dependency-aware program repair via RelationGraph analysis

  • Bib key: tang2025synfix
  • Year: 2025
  • Authors: Tang, Xunzhu; Gao, Jiechao; Xu, Jin; Sun, Tiezhu; Song, Yewei; Ezzini, Saad; Ou{\'e}draogo, Wendk{\^u}uni C; Klein, Jacques; Bissyand{\'e}, Tegawend{\'e} F
  • Venue: Findings of the Association for Computational Linguistics: ACL 2025
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation

Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • Bib key: tao2025code
  • Year: 2025
  • Authors: Tao, Hongyuan; Zhang, Ying; Tang, Zhenhao; Peng, Hongen; Zhu, Xukun; Liu, Bingchang; Yang, Yingguang; Zhang, Ziyin; Xu, Zhaogui; Zhang, Haipeng; others
  • Venue: arXiv preprint arXiv:2505.16901
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Retrieval-Augmented Generation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation

Meta-RAG on Large Codebases Using Code Summarization

  • Bib key: tawosi2025meta
  • Year: 2025
  • Authors: Tawosi, Vali; Alamir, Salwa; Liu, Xiaomo; Veloso, Manuela
  • Venue: arXiv preprint arXiv:2508.02611
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: background, findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Retrieval-Augmented Generation

  • RQ5 / Usage Stages: Localization, Patch Generation

CoreThink: A Symbolic Reasoning Layer to reason over Long Horizon Tasks with LLMs

  • Bib key: vaghasiya2025corethink
  • Year: 2025
  • Authors: Vaghasiya, Jay; Ghugarkar, Omkar; Bhat, Vishvesh; Dholaria, Vipul; McAuley, Julian
  • Venue: arXiv preprint arXiv:2509.00971
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation

Repeton: Structured Bug Repair with ReAct-Guided Patch-and-Test Cycles

  • Bib key: vinh2025repeton
  • Year: 2025
  • Authors: Vinh, Nguyen Phu; Hoang, Anh Chung; Ngo, Chris; Hy, Truong-Son
  • Venue: arXiv preprint arXiv:2506.08173
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Validation, Reproduction Test Generation

Aegis: An agent-based framework for bug reproduction from issue descriptions

  • Bib key: wang2025aegis
  • Year: 2025
  • Authors: Wang, Xinchen; Gao, Pengfei; Meng, Xiangxin; Peng, Chao; Hu, Ruida; Lin, Yun; Gao, Cuiyun
  • Venue: Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: background, findings, introduction

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Reproduction Test Generation

Extracting Conceptual Knowledge to Locate Software Issues

  • Bib key: wang2025extracting
  • Year: 2025
  • Authors: Wang, Ying; Mao, Wenjun; Wang, Chong; Zhou, Zhenhao; Zhou, Yicheng; Zhao, Wenyun; Lou, Yiling; Peng, Xin
  • Venue: arXiv preprint arXiv:2509.21427
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text

  • RQ4 / Knowledge Retrieval: Retrieval-Augmented Generation

  • RQ5 / Usage Stages: Localization

Improving Code Localization with Repository Memory

  • Bib key: wang2025improving
  • Year: 2025
  • Authors: Wang, Boshi; Xu, Weijian; Li, Yunsheng; Gao, Mei; Xie, Yujia; Sun, Huan; Chen, Dongdong
  • Venue: arXiv preprint arXiv:2510.01003
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings, introduction

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Unstructured Text

  • RQ4 / Knowledge Retrieval: Tool Invocation

  • RQ5 / Usage Stages: Localization

Mcts-refined cot: High-quality fine-tuning data for llm-based repository issue resolution

  • Bib key: wang2025mcts
  • Year: 2025
  • Authors: Wang, Yibo; Peng, Zhihao; Wang, Ying; Wei, Zhao; Yu, Hai; Zhu, Zhiliang
  • Venue: arXiv preprint arXiv:2506.12728
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Repository History

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation

A Practitioner's Guide to Multi-turn Agentic Reinforcement Learning

  • Bib key: wang2025practitioner
  • Year: 2025
  • Authors: Wang, Ruiyi; Ammanabrolu, Prithviraj
  • Venue: arXiv preprint arXiv:2510.01132
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation

SeamlessFlow: A Trainer Agent Isolation RL Framework Achieving Bubble-Free Pipelines via Tag Scheduling

  • Bib key: wang2025seamlessflow
  • Year: 2025
  • Authors: Wang, Jinghui; Wang, Shaojie; Cui, Yinghan; Chen, Xuxing; Wang, Chao; Zhang, Xiaojiang; Zhang, Minglei; Zhang, Jiarong; Zhuang, Wenhao; Cao, Yuchen; others
  • Venue: arXiv preprint arXiv:2508.11553
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation, Task Planning

SWE-Dev: Building Software Engineering Agents with Training and Inference Scaling

  • Bib key: wang2025swe
  • Year: 2025
  • Authors: Wang, Haoran; Hou, Zhenyu; Wei, Yao; Tang, Jie; Dong, Yuxiao
  • Venue: arXiv preprint arXiv:2506.07636
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Model Training, Patch Generation, Reproduction Test Generation

Swe-rl: Advancing llm reasoning via reinforcement learning on open software evolution

  • Bib key: wei2025swe
  • Year: 2025
  • Authors: Wei, Yuxiang; Duchenne, Olivier; Copet, Jade; Carbonneaux, Quentin; Zhang, Lingming; Fried, Daniel; Synnaeve, Gabriel; Singh, Rishabh; Wang, Sida I
  • Venue: arXiv preprint arXiv:2502.18449
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Reproduction Test Generation

Toward training superintelligent software agents through self-play swe-rl

  • Bib key: wei2025toward
  • Year: 2025
  • Authors: Wei, Yuxiang; Sun, Zhiqing; McMilin, Emily; Gehring, Jonas; Zhang, David; Synnaeve, Gabriel; Fried, Daniel; Zhang, Lingming; Wang, Sida
  • Venue: arXiv preprint arXiv:2512.18552
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation, Reproduction Test Generation

Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases

  • Bib key: wong2025confucius
  • Year: 2025
  • Authors: Wong, Sherman; Qi, Zhenting; Wang, Zhaodong; Hu, Nathan; Lin, Samuel; Ge, Jun; Gao, Erwin; Chen, Wenlin; Du, Yilun; Yu, Minlan; others
  • Venue: arXiv preprint arXiv:2512.10398
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

Git context controller: Manage the context of llm-based agents like git

  • Bib key: wu2025git
  • Year: 2025
  • Authors: Wu, Junde
  • Venue: arXiv preprint arXiv:2508.00031
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Design Architecture Knowledge, Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

Demystifying llm-based software engineering agents

  • Bib key: xia2025demystifying
  • Year: 2025
  • Authors: Xia, Chunqiu Steven; Deng, Yinlin; Dunn, Soren; Zhang, Lingming
  • Venue: Proceedings of the ACM on Software Engineering
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: background, findings, introduction

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation

Live-SWE-agent: Can Software Engineering Agents Self-Evolve on the Fly?

  • Bib key: xia2025live
  • Year: 2025
  • Authors: Xia, Chunqiu Steven; Wang, Zhe; Yang, Yan; Wei, Yuxiang; Zhang, Lingming
  • Venue: arXiv preprint arXiv:2511.13646
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

Improving the efficiency of LLM agent systems through trajectory reduction

  • Bib key: xiao2025improving
  • Year: 2025
  • Authors: Xiao, Yuan-An; Gao, Pengfei; Peng, Chao; Xiong, Yingfei
  • Venue: arXiv preprint arXiv:2509.23586
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation

Swe-fixer: Training open-source llms for effective and efficient github issue resolution

  • Bib key: xie2025swe
  • Year: 2025
  • Authors: Xie, Chengxing; Li, Bowen; Gao, Chang; Du, He; Lam, Wai; Zou, Difan; Chen, Kai
  • Venue: arXiv preprint arXiv:2501.05040
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation

Think-Search-Patch: A Retrieval-Augmented Reasoning Framework for Repository-Level Code Repair

  • Bib key: xiong2025think
  • Year: 2025
  • Authors: Xiong, Bojian; Lei, Yikun; Liu, Xikai; Zhang, Shaowei; Zhu, Pengyun; Liu, Yan; Leng, Yongqi; Shi, Ling; Zhong, Meizhi; Zhang, Yurong; others
  • Venue: Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: Industry Track
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Retrieval-Augmented Generation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation

Enhancing repository-level software repair via repository-aware knowledge graphs

  • Bib key: yang2025enhancing
  • Year: 2025
  • Authors: Yang, Boyang; Tian, Haoye; Ren, Jiadong; Jin, Shunfu; Liu, Yang; Liu, Feng; Le, Bach
  • Venue: arXiv preprint arXiv:2503.21710
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings, introduction

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation

KGCompass: Knowledge Graph Enhanced Repository-Level Software Repair

  • Bib key: yang2025kgcompass
  • Year: 2025
  • Authors: YANG, BOYANG; REN, JIADONG; JIN, SHUNFU; LIU, YANG; LIU, FENG; LE, BACH; TIAN, HAOYE
  • Venue: N/A
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Structured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation

Kimi-dev: Agentless training as skill prior for swe-agents

  • Bib key: yang2025kimi
  • Year: 2025
  • Authors: Yang, Zonghan; Wang, Shengjie; Fu, Kelin; He, Wenyang; Xiong, Weimin; Liu, Yibo; Miao, Yibo; Gao, Bofei; Wang, Yejie; Ma, Yingwei; others
  • Venue: arXiv preprint arXiv:2509.23045
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation, Reproduction Test Generation

Lingxi: Repository-Level Issue Resolution Framework Enhanced by Procedural Knowledge Guided Scaling

  • Bib key: yang2025lingxi
  • Year: 2025
  • Authors: Yang, Xu; Zhou, Jiayuan; Pacheco, Michael; Zhu, Wenhan; He, Pengfei; Wang, Shaowei; Liu, Kui; Pan, Ruiqi
  • Venue: arXiv preprint arXiv:2510.11838
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Design Architecture Knowledge, Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation, Task Planning

Swe-smith: Scaling data for software engineering agents

  • Bib key: yang2025swe
  • Year: 2025
  • Authors: Yang, John; Lieret, Kilian; Jimenez, Carlos E; Wettig, Alexander; Khandpur, Kabir; Zhang, Yanzhe; Hui, Binyuan; Press, Ofir; Schmidt, Ludwig; Yang, Diyi
  • Venue: arXiv preprint arXiv:2504.21798
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert, Repository History

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Environment Building, Localization, Model Training, Patch Generation, Patch Validation, Reproduction Test Generation

Building Coding Agents via Entropy-Enhanced Multi-Turn Preference Optimization

  • Bib key: yu2025building
  • Year: 2025
  • Authors: Yu, Jiahao; Cheng, Zelei; Wu, Xian; Xing, Xinyu
  • Venue: arXiv preprint arXiv:2509.12434
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Document, Dynamic Execution Results

  • RQ2 / Extraction Methods: Manual Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation

Orcaloca: An llm agent framework for software issue localization

  • Bib key: yu2025orcaloca
  • Year: 2025
  • Authors: Yu, Zhongming; Zhang, Hejia; Zhao, Yujie; Huang, Hanxian; Yao, Matrix; Ding, Ke; Zhao, Jishen
  • Venue: arXiv preprint arXiv:2502.00350
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Reproduction Test Generation

Utboost: Rigorous evaluation of coding agents on swe-bench

  • Bib key: yu2025utboost
  • Year: 2025
  • Authors: Yu, Boxi; Zhu, Yuxuan; He, Pinjia; Kang, Daniel
  • Venue: arXiv preprint arXiv:2506.09289
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Patch Validation, Reproduction Test Generation

Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents

  • Bib key: zainullina2025guided
  • Year: 2025
  • Authors: Zainullina, Karina; Golubev, Alexander; Trofimova, Maria; Polezhaev, Sergei; Badertdinov, Ibragim; Litvintseva, Daria; Karasik, Simon; Fisin, Filipp; Skvortsov, Sergei; Nekrashevich, Maksim; others
  • Venue: arXiv preprint arXiv:2505.13652
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation

Satori-SWE: Evolutionary Test-Time Scaling for Sample-Efficient Software Engineering

  • Bib key: zeng2025satori
  • Year: 2025
  • Authors: Zeng, Guangtao; Shen, Maohao; Chen, Delin; Qi, Zhenting; Das, Subhro; Gutfreund, Dan; Cox, David; Wornell, Gregory; Lu, Wei; Hong, Zhang-Wei; others
  • Venue: arXiv preprint arXiv:2505.23604
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation

Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs

  • Bib key: zeng2025skywork
  • Year: 2025
  • Authors: Zeng, Liang; Li, Yongcong; Xiao, Yuzhen; Li, Changshi; Liu, Chris Yuhao; Yan, Rui; Wei, Tianwen; He, Jujie; Song, Xuchen; Liu, Yang; others
  • Venue: arXiv preprint arXiv:2506.19290
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Environment Building, Localization, Model Training, Patch Generation, Patch Validation

cast: Enhancing code retrieval-augmented generation with structural chunking via abstract syntax tree

  • Bib key: zhang2025cast
  • Year: 2025
  • Authors: Zhang, Yilin; Zhao, Xinran; Wang, Zora Zhiruo; Yang, Chenyang; Wei, Jiayi; Wu, Tongshuang
  • Venue: arXiv preprint arXiv:2506.15655
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Retrieval-Augmented Generation

  • RQ5 / Usage Stages: Localization, Patch Generation

Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

  • Bib key: zhang2025darwin
  • Year: 2025
  • Authors: Zhang, Jenny; Hu, Shengran; Lu, Cong; Lange, Robert; Clune, Jeff
  • Venue: arXiv preprint arXiv:2505.22954
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

One Tool Is Enough: Reinforcement Learning for Repository-Level LLM Agents

  • Bib key: zhang2025one
  • Year: 2025
  • Authors: Zhang, Zhaoxi; Duan, Yitong; Zhang, Yanzhi; Xu, Yiming; Wang, Zhixiang; Liang, Kun; Li, Yang; Liang, Jiahui; Xia, Deguo; Huang, Jizhou; others
  • Venue: arXiv preprint arXiv:2512.20957
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training

Sealign: Alignment training for software engineering agent

  • Bib key: zhang2025sealign
  • Year: 2025
  • Authors: Zhang, Kechi; Zhang, Huangzhao; Li, Ge; You, Jinliang; Li, Jia; Zhao, Yunfei; Jin, Zhi
  • Venue: arXiv preprint arXiv:2503.18455
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection

  • RQ5 / Usage Stages: Model Training

Tom-swe: User mental modeling for software engineering agents

  • Bib key: zhou2025tom
  • Year: 2025
  • Authors: Zhou, Xuhui; Chen, Valerie; Wang, Zora Zhiruo; Neubig, Graham; Sap, Maarten; Wang, Xingyao
  • Venue: arXiv preprint arXiv:2510.21903
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Design Architecture Knowledge, Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

Training Versatile Coding Agents in Synthetic Environments

  • Bib key: zhu2025training
  • Year: 2025
  • Authors: Zhu, Yiqi; Gandhi, Apurva; Neubig, Graham
  • Venue: arXiv preprint arXiv:2512.12216
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Design Architecture Knowledge, Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Document, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation, Task Planning

Swe-search: Enhancing software agents with monte carlo tree search and iterative refinement

  • Bib key: antoniades2024swe
  • Year: 2024
  • Authors: Antoniades, Antonis; {\"O}rwall, Albert; Zhang, Kexun; Xie, Yuxi; Goyal, Anirudh; Wang, William
  • Venue: arXiv preprint arXiv:2410.20285
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation, Task Planning

Masai: Modular architecture for software-engineering ai agents

  • Bib key: arora2024masai
  • Year: 2024
  • Authors: Arora, Daman; Sonwane, Atharv; Wadhwa, Nalin; Mehrotra, Abhav; Utpala, Saiteja; Bairi, Ramakrishna; Kanade, Aditya; Natarajan, Nagarajan
  • Venue: arXiv preprint arXiv:2406.11638
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: background, findings, introduction

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation

Large language monkeys: Scaling inference compute with repeated sampling

  • Bib key: brown2024large
  • Year: 2024
  • Authors: Brown, Bradley; Juravsky, Jordan; Ehrlich, Ryan; Clark, Ronald; Le, Quoc V; R{\'e}, Christopher; Mirhoseini, Azalia
  • Venue: arXiv preprint arXiv:2407.21787
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation

Coder: Issue resolving with multi-agent and task graphs

  • Bib key: chen2024coder
  • Year: 2024
  • Authors: Chen, Dong; Lin, Shaoxin; Zeng, Muhan; Zan, Daoguang; Wang, Jian-Gang; Cheshkov, Anton; Sun, Jun; Yu, Hao; Dong, Guoliang; Aliev, Artem; others
  • Venue: arXiv preprint arXiv:2406.01304
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert, Repository History

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation, Task Planning

Exploring the Potential of Conversational Test Suite Based Program Repair on SWE-bench

  • Bib key: cheshkov2024exploring
  • Year: 2024
  • Authors: Cheshkov, Anton; Zadorozhny, Pavel; Levichev, Rodion; Maslov, Evgeny; Jaldin, Ronaldo Franco
  • Venue: arXiv preprint arXiv:2410.04485
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: background, findings, introduction

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Patch Generation, Patch Validation

SuperCoder2. 0: Technical Report on Exploring the feasibility of LLMs as Autonomous Programmer

  • Bib key: gautam2024supercoder2
  • Year: 2024
  • Authors: Gautam, Anmol; Kumar, Kishore; Jha, Adarsh; NS, Mukunda; Bhola, Ishaan
  • Venue: arXiv preprint arXiv:2409.11190
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation

Can github issues be solved with tree of thoughts?

  • Bib key: la2024can
  • Year: 2024
  • Authors: La Rosa, Ricardo; Hulse, Corey; Liu, Bangdi
  • Venue: arXiv preprint arXiv:2405.13057
  • Reference role: context_reference
  • Screening status: excluded
  • Citation sections: none

Infant agent: A tool-integrated, logic-driven agent with cost-effective api usage

  • Bib key: lei2024infant
  • Year: 2024
  • Authors: Lei, Bin; Li, Yuchen; Zeng, Yiming; Ren, Tao; Luo, Yi; Shi, Tianyu; Gao, Zitian; Hu, Zeyu; Kang, Weitai; Chen, Qiuwu
  • Venue: arXiv preprint arXiv:2411.01114
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

Codetree: Agent-guided tree search for code generation with large language models

  • Bib key: li2024codetree
  • Year: 2024
  • Authors: Li, Jierui; Le, Hung; Zhou, Yingbo; Xiong, Caiming; Savarese, Silvio; Sahoo, Doyen
  • Venue: arXiv preprint arXiv:2411.04329
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase

  • RQ2 / Extraction Methods: Model-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Retrieval-Augmented Generation

  • RQ5 / Usage Stages: Localization, Patch Generation

Llms as continuous learners: Improving the reproduction of defective code in software issues

  • Bib key: lin2024llms
  • Year: 2024
  • Authors: Lin, Yalan; Ma, Yingwei; Cao, Rongyu; Li, Binhua; Huang, Fei; Gu, Xiaodong; Li, Yongbin
  • Venue: arXiv preprint arXiv:2411.13941
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Reproduction Test Generation

Codexembed: A generalist embedding model family for multiligual and multi-task code retrieval

  • Bib key: liu2024codexembed
  • Year: 2024
  • Authors: Liu, Ye; Meng, Rui; Joty, Shafiq; Savarese, Silvio; Xiong, Caiming; Zhou, Yingbo; Yavuz, Semih
  • Venue: arXiv preprint arXiv:2411.12644
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Retrieval-Augmented Generation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation

Codexgraph: Bridging large language models and code repositories via code graph databases

  • Bib key: liu2024codexgraph
  • Year: 2024
  • Authors: Liu, Xiangyan; Lan, Bo; Hu, Zhiyuan; Liu, Yang; Zhang, Zhicheng; Wang, Fei; Shieh, Michael; Zhou, Wenmeng
  • Venue: arXiv preprint arXiv:2408.03910
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Graph

  • RQ4 / Knowledge Retrieval: Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation

Marscode agent: Ai-native automated bug fixing

  • Bib key: liu2024marscode
  • Year: 2024
  • Authors: Liu, Yizhou; Gao, Pengfei; Wang, Xinchen; Liu, Jie; Shi, Yexuan; Zhang, Zhao; Peng, Chao
  • Venue: arXiv preprint arXiv:2409.00899
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation

Lingma swe-gpt: An open development-process-centric language model for automated software improvement

  • Bib key: ma2024lingma
  • Year: 2024
  • Authors: Ma, Yingwei; Cao, Rongyu; Cao, Yongchang; Zhang, Yue; Chen, Jue; Liu, Yibo; Liu, Yuchen; Li, Binhua; Huang, Fei; Li, Yongbin
  • Venue: arXiv preprint arXiv:2411.00622
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Localization, Model Training, Patch Generation, Patch Validation, Task Planning

Repository Structure-Aware Training Makes SLMs Better Issue Resolver

  • Bib key: ma2024repository
  • Year: 2024
  • Authors: Ma, Zexiong; An, Shengnan; Lin, Zeqi; Zou, Yanzhen; Xie, Bing
  • Venue: arXiv preprint arXiv:2412.19031
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Repository History

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection

  • RQ5 / Usage Stages: Model Training

Repograph: Enhancing ai software engineering with repository-level code graph

  • Bib key: ouyang2024repograph
  • Year: 2024
  • Authors: Ouyang, Siru; Yu, Wenhao; Ma, Kaixin; Xiao, Zilin; Zhang, Zhihan; Jia, Mengzhao; Han, Jiawei; Zhang, Hongming; Yu, Dong
  • Venue: arXiv preprint arXiv:2410.14684
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: background, findings, introduction

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Graph, Unstructured Text

  • RQ4 / Knowledge Retrieval: Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation

Training software engineering agents and verifiers with swe-gym

  • Bib key: pan2024training
  • Year: 2024
  • Authors: Pan, Jiayi; Wang, Xingyao; Neubig, Graham; Jaitly, Navdeep; Ji, Heng; Suhr, Alane; Zhang, Yizhe
  • Venue: arXiv preprint arXiv:2412.21139
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Background Knowledge, Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge, Technology Stack Knowledge

  • RQ2 / Knowledge Sources: Codebase, Document, Dynamic Execution Results

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Tool Invocation

  • RQ5 / Usage Stages: Environment Building, Localization, Model Training, Patch Generation, Patch Validation

Hyperagent: Generalist software engineering agents to solve coding tasks at scale

  • Bib key: phan2024hyperagent
  • Year: 2024
  • Authors: Phan, Huy Nhat; Nguyen, Tien N; Nguyen, Phong X; Bui, Nghi DQ
  • Venue: arXiv preprint arXiv:2409.16299
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation, Task Planning

Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents

  • Bib key: rombaut2024watson
  • Year: 2024
  • Authors: Rombaut, Benjamin; Masoumzadeh, Sogol; Vasilevski, Kirill; Lin, Dayi; Hassan, Ahmed E
  • Venue: arXiv preprint arXiv:2411.03455
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Model-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt

  • RQ5 / Usage Stages: Localization

CoRNStack: High-quality contrastive data for better code retrieval and reranking

  • Bib key: suresh2024cornstack
  • Year: 2024
  • Authors: Suresh, Tarun; Reddy, Revanth Gangi; Xu, Yifei; Nussbaum, Zach; Mulyar, Andriy; Duderstadt, Brandon; Ji, Heng
  • Venue: arXiv preprint arXiv:2412.01007
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Parametric Injection, Retrieval-Augmented Generation

  • RQ5 / Usage Stages: Localization, Model Training

Magis: Llm-based multi-agent framework for github issue resolution

  • Bib key: tao2024magis
  • Year: 2024
  • Authors: Tao, Wei; Zhou, Yucheng; Wang, Yanlin; Zhang, Wenqiang; Zhang, Hongyu; Cheng, Yu
  • Venue: Advances in Neural Information Processing Systems
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: background, findings, introduction

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Repository Evolution Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Repository History

  • RQ2 / Extraction Methods: Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Retrieval-Augmented Generation, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

Openhands: An open platform for ai software developers as generalist agents

  • Bib key: wang2024openhands
  • Year: 2024
  • Authors: Wang, Xingyao; Li, Boxuan; Song, Yufan; Xu, Frank F; Tang, Xiangru; Zhuge, Mingchen; Pan, Jiayi; Song, Yueqi; Li, Bowen; Singh, Jaskirat; others
  • Venue: arXiv preprint arXiv:2407.16741
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: background, findings, introduction

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Structured Text, Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Task Planning

Swe-agent: Agent-computer interfaces enable automated software engineering

  • Bib key: yang2024swe
  • Year: 2024
  • Authors: Yang, John; Jimenez, Carlos E; Wettig, Alexander; Lieret, Kilian; Yao, Shunyu; Narasimhan, Karthik; Press, Ofir
  • Venue: Advances in Neural Information Processing Systems
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: background, findings, introduction

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation, Reproduction Test Generation, Task Planning

Autocoderover: Autonomous program improvement

  • Bib key: zhang2024autocoderover
  • Year: 2024
  • Authors: Zhang, Yuntong; Ruan, Haifeng; Fan, Zhiyu; Roychoudhury, Abhik
  • Venue: Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: background, findings, introduction

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Development Tool Knowledge, Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase, Dynamic Execution Results

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt, Tool Invocation

  • RQ5 / Usage Stages: Localization, Patch Generation, Patch Validation

CodeV: Issue Resolving with Visual Data

  • Bib key: zhang2024codev
  • Year: 2024
  • Authors: Zhang, Linhao; Zan, Daoguang; Yang, Quanshun; Huang, Zhirong; Chen, Dong; Shen, Bo; Liu, Tianyu; Gong, Yongshun; Huang, Pengjie; Lu, Xudong; others
  • Venue: arXiv preprint arXiv:2412.17315
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: background, findings

  • RQ1 / Knowledge Layers: Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge

  • RQ2 / Knowledge Sources: Codebase

  • RQ2 / Extraction Methods: Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt

  • RQ5 / Usage Stages: Localization, Patch Generation

Diversity empowers intelligence: Integrating expertise of software engineering agents

  • Bib key: zhang2024diversity
  • Year: 2024
  • Authors: Zhang, Kexun; Yao, Weiran; Liu, Zuxin; Feng, Yihao; Liu, Zhiwei; Murthy, Rithesh; Lan, Tian; Li, Lei; Lou, Renze; Xu, Jiacheng; others
  • Venue: arXiv preprint arXiv:2408.07060
  • Reference role: core_corpus
  • Screening status: included
  • Citation sections: findings

  • RQ1 / Knowledge Layers: Procedural Knowledge, Repository Knowledge

  • RQ1 / Knowledge Types: Existing Code Knowledge, Experiential Knowledge

  • RQ2 / Knowledge Sources: Codebase, Expert

  • RQ2 / Extraction Methods: Manual Extraction, Model-based Extraction, Rule-based Extraction

  • RQ3 / Representation Formats: Unstructured Text

  • RQ4 / Knowledge Retrieval: Direct Prompt

  • RQ5 / Usage Stages: Patch Validation

Swe-bench: Can language models resolve real-world github issues?

  • Bib key: jimenez2023swe
  • Year: 2023
  • Authors: Jimenez, Carlos E; Yang, John; Wettig, Alexander; Yao, Shunyu; Pei, Kexin; Press, Ofir; Narasimhan, Karthik
  • Venue: arXiv preprint arXiv:2310.06770
  • Reference role: benchmark_background
  • Screening status: seed
  • Citation sections: background, introduction, methodology