Benchmark results
Open-ended code and Galaxy-API code performance on computational biology tasks.
Benchmarks
Selected benchmark results
Each selected benchmark is shown separately.
Every bar is a benchmark-specific result; no cross-benchmark average or overall ranking is calculated.
BixBench Verified-50
50 tasks · 3 runs per model and environment
Overall accuracy: Open-ended code 87.2% · Galaxy-API code 87.7%
BixBench examples
Open a task to see each model's three Open-ended code and Galaxy-API code runs. ✓ pass · × fail
Transcriptomicsbix-30-q3Multiple-testing correctionCompare the number of significant miRNAs after Bonferroni and Benjamini–Yekutieli correction.
67% passView results
Compare the number of significant miRNAs after Bonferroni and Benjamini–Yekutieli correction.
Phylogeneticsbix-45-q1RCV score comparisonUse PhyKIT and a Mann–Whitney U test to compare animal and fungal orthologs.
33% passView results
Use PhyKIT and a Mann–Whitney U test to compare animal and fungal orthologs.
Genomicsbix-55-q1BUSCO completenessFind the single-copy orthologs that are complete and present in all four proteomes.
92% passView results
Find the single-copy orthologs that are complete and present in all four proteomes.
CompBioBench
100 tasks · 24 published replicates
Three-replicate means. GPT-5.5 Galaxy-API code uses official leaderboard scores for R1 (84%), R2 (89%), and R3 (87%). GPT-5.6 Sol Galaxy-API code uses official scores for R1 (93%), R2 (91%), and R3 (92%). GPT-5.6 Luna uses official scores for Open-ended code R1 (84%), R2 (86%), and R3 (85%), and Galaxy-API code R1 (86%), R2 (84%), and R3 (85%). DeepSeek-v4-pro-0813 uses official scores for Open-ended code R1 (80%), R2 (87%), and R3 (86%), and Galaxy-API code R1 (84%), R2 (87%), and R3 (83%).
CompBioBench examples
Ground truth is unavailable. These marks show our predicted outcomes: ✓ predicted correct · × predicted incorrect
Machine learningborzoi-rnaseq-q1Borzoi RNA-seq predictionRecover total forward-strand RNA-seq coverage from a Borzoi model prediction over a genomic interval.
92% predicted correctView predictions
Recover total forward-strand RNA-seq coverage from a Borzoi model prediction over a genomic interval.
Single-cellthree-way-barnyard-q1Three-species mixtureEstimate human, mouse, and pig proportions from paired-end 10x scRNA-seq reads.
100% predicted correctView predictions
Estimate human, mouse, and pig proportions from paired-end 10x scRNA-seq reads.
Single-cellthree-way-barnyard-q2Mixture and tissue originRecover species proportions and assign a tissue of origin for each species.
67% predicted correctView predictions
Recover species proportions and assign a tissue of origin for each species.
IWC Workflow Benchmark
10 end-to-end workflow tasks · 3 runs per model and environment
Overall mean final-output score: Open-ended code 94.6% · Galaxy-API code 98.2%. A score of 100% is a perfect match to the benchmark output.
IWC Workflow Benchmark examples
Each row shows R1, R2, and R3 as percentages.
Microbiomewf00517-sample 16S ASV tableBuild an exact sequence-variant count table from paired 16S reads while preserving all 17 case-sensitive sample identifiers. Two Open-ended code runs lowercased every identifier and received zero. We score both the recovered ASV sequences and their per-sample raw counts, then combine the two F1 scores using a geometric mean. IWC workflow ↗
Open-ended 79.7%Galaxy-API 95.9%
Build an exact sequence-variant count table from paired 16S reads while preserving all 17 case-sensitive sample identifiers. Two Open-ended code runs lowercased every identifier and received zero. We score both the recovered ASV sequences and their per-sample raw counts, then combine the two F1 scores using a geometric mean. IWC workflow ↗
Genome assemblywf007Mitochondrial genome from PacBio HiFi readsAssemble one complete mitochondrial sequence for Agrius convolvuli. Three runs selected the wrong assembly candidate; most other non-perfect sequences differed from the benchmark at only a small number of bases. We score matching canonical 31-base sequence pieces with F1, so reverse-complement orientation is accepted while missing, extra, or duplicated sequence lowers the score. IWC workflow ↗
Open-ended 82.6%Galaxy-API 91.3%
Assemble one complete mitochondrial sequence for Agrius convolvuli. Three runs selected the wrong assembly candidate; most other non-perfect sequences differed from the benchmark at only a small number of bases. We score matching canonical 31-base sequence pieces with F1, so reverse-complement orientation is accepted while missing, extra, or duplicated sequence lowers the score. IWC workflow ↗
Metaproteomicswf009Novel-peptide validationValidate microbial candidate peptides against tandem mass spectra and report every supported peptide-to-protein pair. Most score loss came from accepting extra pairs, rather than missing the target pairs. We compute F1 over the peptide–protein pairs after removing peptide modification labels and normalizing protein identifiers to UniProt accessions. IWC workflow ↗
Open-ended 93.9%Galaxy-API 95.8%
Validate microbial candidate peptides against tandem mass spectra and report every supported peptide-to-protein pair. Most score loss came from accepting extra pairs, rather than missing the target pairs. We compute F1 over the peptide–protein pairs after removing peptide modification labels and normalizing protein identifiers to UniProt accessions. IWC workflow ↗
Select at least one benchmark.