aisa-group/PostTrainBench
538 stars · Last commit 2026-08-21
Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 10 hours
README preview
# PostTrainBench: Can LLM Agents Automate LLM Post-Training? [](http://posttrainbench.com/) We introduce PostTrainBench, a benchmark that measures the ability of CLI agents to post-train pre-trained large language models (LLMs). In PostTrainBench, the agent's task is to improve the performance of a base LLM on a given benchmark. The agent is given access to an evaluation script and 10 hours on an H100 GPU. Performance is measured by the benchmark score of the post-trained LLM. This setup naturally evaluates an agent's ability to conduct AI R&D. > [!IMPORTANT] > **Harbor support coming soon!** This repository currently targets our internal HPC cluster (HTCondor). We are adding [Harbor](https://github.com/harbor-framework/harbor) support to make it straightforward to run on rented hardware (e.g., cloud GPUs). See our [PR](https://github.com/aisa-group/PostTrainBench/pull/8). ## Scaffolds Agents are run through one of 4 CLI scaffolds: Claude Code, Codex CLI, Gemini CLI, and OpenCode. ## Evaluation Tasks PostTrainBench includes 7 benchmarks spanning reasoning, tool use, knowledge, math, health, and code: 1. **AIME 2025** — Math competition problems 2. **Arena Hard Writing** — Creative writing benchmark adapted from ArenaHard v2