DSPy GEPA vs MiPRO Benchmark & Optimization Platform

Full-corpus optimization results across ALL 176 taxonomy fields on 100 legal agreements comparing DSPy GEPA vs MiPRO.

Compile Models: 5.6-luna (Task) · 5.6-sol (Reflection)
Overall Winner: DSPy GEPA Did MiPRO beat GEPA? No.

GEPA beat MiPRO by +0.132 F1 corpus average (0.842 F1 vs 0.710 F1), winning 112 out of 140 evaluatable fields (80.0%). MiPRO won/tied on 28 fields (20.0%) where Bayesian seed search found top exemplars at $0 reflection cost, but plateaued early on complex clauses.

GEPA Mean F1
0.842
MiPRO Mean F1
0.710
Head-to-Head
112 to 28
Paired F1 Delta
+63.7%
Display Optimizer Trajectories (All 176 Fields):
Cohort Filters:

10-Iteration Optimization Curves across ALL 176 Taxonomy Fields

Showing pale purple curves for all 176 GEPA runs and pale orange curves for all 176 MiPRO runs, with bold corpus averages on top.

Bold GEPA Mean (0.842 F1)
Bold MiPRO Mean (0.710 F1)
176 Fields (GEPA: Pale Purple)
176 Fields (MiPRO: Pale Orange)

Agreement Type

Agreement Metadata Tuned GEPA Winner (+0.18 F1)
Domain: Agreement Metadata · Status: Tuned · GEPA Final: 0.88 (+63%) · MiPRO Final: 0.70
GEPA F1
0.54 → 0.88 (+63%)
MiPRO F1
0.54 → 0.70 (+30%)
Head-to-Head Delta
GEPA +0.18
Baseline Zero-Shot Prompt:
Tuned Instruction via DSPy GEPA (GPT-5.6 Sol Reflection):

All 176 Field Outcomes & Winner Breakdown

Check any row box to highlight that field's purple GEPA and orange MiPRO trajectory on the chart above.

Tick Field Name Domain Winner GEPA F1 MiPRO F1 Base (0) Iter 1 Iter 2 Iter 3 Iter 4 Iter 5 Iter 6 Iter 7 Iter 8 Iter 9 Iter 10 Tokens

Actual Real Costs & Token Telemetry (176 Fields across 100 Documents)

Standardized against 0% prompt cache hit rate (Gross List Price baseline). Eliminates test rerun cache artifacts for true production modeling.

100% Run Tokens Counted Standard 0% Cache Baseline
DSPy GEPA Real Cost & Token Telemetry 0.842 F1 (Winner)
Total Gross Run Tokens: 294.3M tokens (All candidate + reflection tokens)
Standard List-Price Compute (0% Cache): $194.92 (~$1.107 / field)
Candidate Evaluation (5.6-luna): 289.5M tokens ($106.14 gross list price)
GPT-5.6 Sol Reflection Mutations: 4.8M tokens ($88.78 @ $5 in / $30 out / 2.0x CoT)
Steady-State Pipeline (40% Cache Hits): $122.71 (~$0.697 / field)
Benchmark Test Rerun (95% Local Hits): $14.43 (Repetitive test rerun artifact)
Mean Compilation Time: 240.2s / field (+63.7% accuracy jump)
DSPy MiPRO Real Cost & Token Telemetry 0.710 F1 (Baseline)
Total Gross Run Tokens: 301.1M tokens (Candidate bootstrap permutations)
Standard List-Price Compute (0% Cache): $131.33 (~$0.746 / field)
Candidate Evaluation (5.6-luna): 301.1M tokens ($131.33 gross list price)
Reflection Mutations: 0 tokens ($0.00 reflection compute)
Steady-State Pipeline (40% Cache Hits): $78.80 (~$0.448 / field)
Benchmark Test Rerun (95% Local Hits): $4.22 (Repetitive test rerun artifact)
Mean Compilation Time: 64.8s / field (Fast baseline bootstrapping)
Copied blob JSON to clipboard!