Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

Harness-Aware Evaluation of LLM Agents

Systems, Benchmarks, and Protocols

Project website · arXiv: XXX

Contact: Shaina Raza

LLM agent results depend on more than a backbone model. This survey treats each reported score as the outcome of a complete evaluation configuration:

Model (M) + Harness (H) + Environment (E) + Evaluator (V), tested under an explicit evaluation protocol.

The framework helps distinguish model capability from the effects of context and memory, tool interfaces, execution control, coordination, verification, permissions, environmental conditions, scoring criteria, and resource budgets.

Cited literature

The list includes every work actively cited in the manuscript. It excludes bibliography entries referenced only in commented LaTeX.

Foundations and agent architectures

Agent systems and operational harnesses

Harness design, comparison, and optimization

Context, memory, and procedural knowledge

Tool use, planning, coordination, and interaction

Safety, robustness, and human oversight

Environments and agent benchmarks

Evaluators and evaluation protocols

Related surveys

  • A Survey on Large Language Model Based Autonomous Agents (2024)
  • The Rise and Potential of Large Language Model Based Agents: A Survey (2025)
  • Evaluation and Benchmarking of LLM Agents: A Survey (2025)
  • A Survey on Trustworthy LLM Agents: Threats and Countermeasures (2025)
  • A Survey on the Optimization of Large Language Model-Based Agents (2026)
  • Evaluating LLM-Based Agents for Multi-Turn Conversations: A Survey (2026)
  • Generalizability of Large Language Model-Based Agents: A Comprehensive Survey (2026)
  • A Survey on Evaluation of LLM-Based Agents (2026)
  • Agent Harness Engineering: A Survey (2026)

Citation

@article{radwan2026harness,
  title   = {Harness-Aware Evaluation of LLM Agents:
             Systems, Benchmarks, and Protocols},
  author  = {Radwan, Ahmed Y. and Vasilakos, Athanasios V.
             and Raza, Shaina},
  journal = {arXiv preprint arXiv:XXXX.XXXXX},
  year    = {2026}
}

Acknowledgements

Resources used in preparing this research were provided, in part, by the Province of Ontario, the Government of Canada through CIFAR, and companies sponsoring the Vector Institute. This research was funded by the European Union’s Horizon Europe AIXPERT project (Grant Agreement No. 101214389).

About

A survey of LLM agent evaluation that treats each score as a complete configuration: model, harness, environment, and evaluator.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors