Base model vs trained model, same tests.
The pipeline compares Qwen2.5-Coder-3B base and AxiomCode-3B on the same verified programming tasks with the same runner.
Executable verification for coding models
The pipeline compares Qwen2.5-Coder-3B base and AxiomCode-3B on the same verified programming tasks with the same runner.
The first suite records patch apply rate, pass@1, test pass rate, runtime, and failure category. Results are TBD until the runs exist.
Training examples come from verified coding-task traces: task, files, command output, unified diff, verification result, and split metadata.
The first report compares Qwen2.5-Coder-3B base against AxiomCode-3B. No score is published before the benchmark runs.
The repo is set up for local smoke checks first, then Modal import checks, serving, base evaluation, LoRA training, trained evaluation, and compare reports.
Compute support helps scale the first measured training run. The public identity stays the same: CodeAxiom builds AxiomCode through executable verification.
contact@avixosec.xyz