Horizon Bench contains reproductions of AI Scientist projects.
- Provide easy-to-run, reproducible implementations of various AI Scientist projects. Research code is rarely directly runnable, and often requires hours of debugging to get it to work.
- Provide a common framework (same LLMs, same evaluation method) for comparing different AI Scientist projects.
- Provide a shared code backbone, to enable quick reimplementation of future papers on AI Scientists.
conda env create --name horizon_bench python=3.13 -c conda-forge
conda activate horizon_bench
pip install -r requirements.txt
Pandoc is used for COI Agent's paper text extraction
sudo apt install pandoc
GROBID is used via scipdf-parser to extract full text from OpenReview PDFs.
Requires Java 17 (GROBID's Gradle wrapper does not support Java 21+).
- Install Java 17 and scipdf-parser:
sudo apt install -y openjdk-17-jdk pip install scipdf-parser
- Install GROBID (check which version is currently newest, 0.8.2 at time of writing):
wget https://github.com/kermitt2/grobid/archive/0.8.2.zip unzip 0.8.2.zip cd grobid-0.8.2 ./gradlew clean install - Run GROBID before using the PDF extraction:
GROBID will start on
cd grobid-0.8.2 ./gradlew runhttp://localhost:8070by default.
Replace the workspace path in src/horizon_bench/Config.py according to your setup.
I used absolute paths to avoid conflicts.
jupyter notebook src/ai_researcher/reproducing_ai_researcher.ipynb
Word of caution: AI Researcher enables arbitrary code execution on your machine by the LLM. As far as I can tell it didn't do anything harmful yet. If you want to be extra safe, run my code in a VM.
jupyter notebook src/ai_scientist/reproducing_ai_scientist.ipynb
I reimplemented AI Scientist's LLM as a judge functionality, which aggregates multiple reviews from an LLM into a final review.
src/ai_scientist/llm_as_a_judge/llm_as_a_judge.py
I list here the simplifications that I made for the different projects. For example, I often found asyncio code which was completely redundant, because people simply always called await on each function. At that point, there is no concurrency benefit at all (AI Researcher, COI Agent).
COI Agent uses SciPDF and grobid to extract text from PDFs. This is general, as every paper has a PDF.
However, grobid produces a lot of noisy text and some errors. For now, I replaced this with:
- Download latex source from arxiv
- Convert tex to txt using pandoc This produces much cleaner text. The downside is that this only works for arxiv papers.