Galaxy Agent BenchmarksGitHub ↗

Benchmark results

Open-ended code and Galaxy-API code performance on computational biology tasks.

Benchmarks

Select benchmarks

Selected benchmark results

Each selected benchmark is shown separately.

Open-ended codeGalaxy-API code
GPT-5.5Codex · high reasoning
GPT-5.6 SolCodex · high reasoning
GPT-5.6 LunaCodex · maximum reasoning
DeepSeek-v4-pro-0813Codex · high reasoning

Every bar is a benchmark-specific result; no cross-benchmark average or overall ranking is calculated.

BixBench Verified-50

50 tasks · 3 runs per model and environment

Open-ended codeGalaxy-API code
GPT-5.5Codex · high reasoning
90.0%
91.3%
GPT-5.6 SolCodex · high reasoning
89.3%
90.7%
GPT-5.6 LunaCodex · maximum reasoning
88.7%
86.7%
DeepSeek-v4-pro-0813Codex · high reasoning
80.7%
82.0%

Overall accuracy: Open-ended code 87.2% · Galaxy-API code 87.7%

BixBench examples

Open a task to see each model's three Open-ended code and Galaxy-API code runs. ✓ pass · × fail

Transcriptomicsbix-30-q3
Multiple-testing correction

Compare the number of significant miRNAs after Bonferroni and Benjamini–Yekutieli correction.

67% passView results
GPT-5.5Open-ended code✓✓×Galaxy-API code✓✓✓
GPT-5.6 SolOpen-ended code×××Galaxy-API code✓✓✓
GPT-5.6 LunaOpen-ended code✓✓×Galaxy-API code✓✓✓
DeepSeek-v4-pro-0813Open-ended code×××Galaxy-API code✓✓✓
Phylogeneticsbix-45-q1
RCV score comparison

Use PhyKIT and a Mann–Whitney U test to compare animal and fungal orthologs.

33% passView results
GPT-5.5Open-ended code×××Galaxy-API code×××
GPT-5.6 SolOpen-ended code✓✓✓Galaxy-API code×××
GPT-5.6 LunaOpen-ended code✓×✓Galaxy-API code×××
DeepSeek-v4-pro-0813Open-ended code✓✓✓Galaxy-API code×××
Genomicsbix-55-q1
BUSCO completeness

Find the single-copy orthologs that are complete and present in all four proteomes.

92% passView results
GPT-5.5Open-ended code✓✓✓Galaxy-API code✓✓✓
GPT-5.6 SolOpen-ended code×✓✓Galaxy-API code✓✓✓
GPT-5.6 LunaOpen-ended code✓✓×Galaxy-API code✓✓✓
DeepSeek-v4-pro-0813Open-ended code✓✓✓Galaxy-API code✓✓✓

CompBioBench

100 tasks · 24 published replicates

Open-ended codeGalaxy-API code
GPT-5.6 SolCodex · high reasoning
91.0%
92.0%
GPT-5.5Codex · high reasoning
86.3%
86.7%
DeepSeek-v4-pro-0813Codex · high reasoning
84.3%
84.7%
GPT-5.6 LunaCodex · maximum reasoning
85.0%
85.0%

Three-replicate means. GPT-5.5 Galaxy-API code uses official leaderboard scores for R1 (84%), R2 (89%), and R3 (87%). GPT-5.6 Sol Galaxy-API code uses official scores for R1 (93%), R2 (91%), and R3 (92%). GPT-5.6 Luna uses official scores for Open-ended code R1 (84%), R2 (86%), and R3 (85%), and Galaxy-API code R1 (86%), R2 (84%), and R3 (85%). DeepSeek-v4-pro-0813 uses official scores for Open-ended code R1 (80%), R2 (87%), and R3 (86%), and Galaxy-API code R1 (84%), R2 (87%), and R3 (83%).

CompBioBench examples

Ground truth is unavailable. These marks show our predicted outcomes: ✓ predicted correct · × predicted incorrect

Machine learningborzoi-rnaseq-q1
Borzoi RNA-seq prediction

Recover total forward-strand RNA-seq coverage from a Borzoi model prediction over a genomic interval.

92% predicted correctView predictions
GPT-5.5Open-ended code✓✓×Galaxy-API code✓✓✓
GPT-5.6 SolOpen-ended code✓✓✓Galaxy-API code✓✓✓
DeepSeek-v4-pro-0813Open-ended code×✓✓Galaxy-API code✓✓✓
GPT-5.6 LunaOpen-ended code✓✓✓Galaxy-API code✓✓✓
Single-cellthree-way-barnyard-q1
Three-species mixture

Estimate human, mouse, and pig proportions from paired-end 10x scRNA-seq reads.

100% predicted correctView predictions
GPT-5.5Open-ended code✓✓✓Galaxy-API code✓✓✓
GPT-5.6 SolOpen-ended code✓✓✓Galaxy-API code✓✓✓
DeepSeek-v4-pro-0813Open-ended code✓✓✓Galaxy-API code✓✓✓
GPT-5.6 LunaOpen-ended code✓✓✓Galaxy-API code✓✓✓
Single-cellthree-way-barnyard-q2
Mixture and tissue origin

Recover species proportions and assign a tissue of origin for each species.

67% predicted correctView predictions
GPT-5.5Open-ended code✓✓×Galaxy-API code✓✓✓
GPT-5.6 SolOpen-ended code✓×✓Galaxy-API code✓✓✓
DeepSeek-v4-pro-0813Open-ended code✓×✓Galaxy-API code×✓×
GPT-5.6 LunaOpen-ended code✓✓×Galaxy-API code××✓

IWC Workflow Benchmark

10 end-to-end workflow tasks · 3 runs per model and environment

Open-ended codeGalaxy-API code
GPT-5.5Codex · high reasoning
88.1%
95.4%
GPT-5.6 SolCodex · high reasoning
98.9%
99.1%
GPT-5.6 LunaCodex · maximum reasoning
98.4%
98.5%
DeepSeek-v4-pro-0813Codex · high reasoning
93.0%
99.9%

Overall mean final-output score: Open-ended code 94.6% · Galaxy-API code 98.2%. A score of 100% is a perfect match to the benchmark output.

IWC Workflow Benchmark examples

Each row shows R1, R2, and R3 as percentages.

Microbiomewf005
17-sample 16S ASV table

Build an exact sequence-variant count table from paired 16S reads while preserving all 17 case-sensitive sample identifiers. Two Open-ended code runs lowercased every identifier and received zero. We score both the recovered ASV sequences and their per-sample raw counts, then combine the two F1 scores using a geometric mean. IWC workflow ↗

Open-ended 79.7%Galaxy-API 95.9%
GPT-5.5Open-ended code0.090.00.0Galaxy-API code91.891.895.6
GPT-5.6 SolOpen-ended code95.295.294.9Galaxy-API code95.695.695.1
GPT-5.6 LunaOpen-ended code95.510091.8Galaxy-API code93.895.695.6
DeepSeek-v4-pro-0813Open-ended code93.5100100Galaxy-API code100100100
Genome assemblywf007
Mitochondrial genome from PacBio HiFi reads

Assemble one complete mitochondrial sequence for Agrius convolvuli. Three runs selected the wrong assembly candidate; most other non-perfect sequences differed from the benchmark at only a small number of bases. We score matching canonical 31-base sequence pieces with F1, so reverse-complement orientation is accepted while missing, extra, or duplicated sequence lowers the score. IWC workflow ↗

Open-ended 82.6%Galaxy-API 91.3%
GPT-5.5Open-ended code0.094.299.6Galaxy-API code0.010099.4
GPT-5.6 SolOpen-ended code99.699.699.6Galaxy-API code99.699.6100
GPT-5.6 LunaOpen-ended code99.699.699.6Galaxy-API code99.599.599.5
DeepSeek-v4-pro-0813Open-ended code99.60.099.5Galaxy-API code10099.599.1
Metaproteomicswf009
Novel-peptide validation

Validate microbial candidate peptides against tandem mass spectra and report every supported peptide-to-protein pair. Most score loss came from accepting extra pairs, rather than missing the target pairs. We compute F1 over the peptide–protein pairs after removing peptide modification labels and normalizing protein identifiers to UniProt accessions. IWC workflow ↗

Open-ended 93.9%Galaxy-API 95.8%
GPT-5.5Open-ended code98.284.185.2Galaxy-API code98.290.898.2
GPT-5.6 SolOpen-ended code10094.998.2Galaxy-API code98.298.290.8
GPT-5.6 LunaOpen-ended code91.596.479.6Galaxy-API code98.286.890.0
DeepSeek-v4-pro-0813Open-ended code10099.3100Galaxy-API code100100100

Select at least one benchmark.