---
title: Ora Performance
section: Ora — Performance
status: published
subtitle: What the harness adds, measured.
description: The same analysis, run six ways and scored for quality and cost across 196 of Ora's techniques by an independent judge panel. Capability turns out to come from the harness, not the model.
authors:
  - The Ora Foundation
downloads:
  md: /papers/white/ora-performance.md
license: https://creativecommons.org/publicdomain/zero/1.0/
wide: true
---

# Ora Performance

Ora's harness provides the missing systems layer for AI: structured execution, adversarial review, error correction, drift control, and reruns around interchangeable models. It brings frontier-grade quality to small, locally run models, magnifies frontier-model abilities, and makes complex, multi-step cognitive automation reliable enough to fulfill the long-awaited promise of AI agents.

- **Capability:** Ora makes interchangeable models execute structured analytical techniques and guided cognitive operations accurately, turning models from answer generators into working components inside a larger system.
- **Reliability:** Ora reduces drift and prevents failure through adversarial review, fact checks, course corrections, and reruns when an operation falls below standard.
- **Automation:** Ora makes multi-step AI work viable because its reliability architecture catches bad intermediate steps before errors compound.
- **Local power:** Ora lets small private models reach raw-frontier performance on-device, reducing or eliminating cloud dependency.
- **Economics:** Ora requires fewer user prompts to reach higher-quality answers while allowing more cost-effective models, reducing total costs.

## The results

| Configuration | Rating /10 | Failure rate | Cost / 100 queries | Avg. time / run (min) |
|---|--:|--:|--:|--:|
| **Flagship Model** — frontier model, raw | 5.4 | 24% | $8 | 1.0 |
| **Small Model, Raw** — 9B, no harness | 4.2 | 41% | $0.06 | 1.1 |
| **Small Local Only** — 9B in the harness | 5.4 | 8% | $3 | 18.9 |
| **Optimum** — Ora's recommended setup | 7.6 | 1% | $13 | 22.3 |
| **Optimum + Premium** | 8.4 | 1% | $40 | 23.6 |
| **Premium** — maximum capability | 8.6 | 2% | $228 | 17.3 |

*Ratings are 0–10. Fidelity measures whether the answer actually performed the assigned analysis. Overall is capped at Fidelity, so a fluent answer that does not execute the method cannot score high. Cost is the API-equivalent price per 100 queries; time is wall-clock minutes per run.*

- **Private 9B reaches flagship quality** — a harnessed 9B now matches the raw flagship, making frontier-grade work possible on today's 16GB local machines with complete data privacy.
- **The harness lifts small models** — the same 9B improves ≈30%, from 4.2 raw to 5.4 harnessed.
- **The harness lifts frontier models too** — the raw flagship rises ≈60%, from 5.4 raw to 8.6 inside the Premium harness.
- **Low-cost models beat raw flagship quality** — Optimum improves on the raw flagship by ≈40%, with a 1% failure rate, using inexpensive models and adversarial review.
- **Premium works best as final review** — Optimum + Premium gets most of the benefit by using frontier intelligence only at consolidation.
- **Premium everywhere wastes money** — Optimum + Premium nearly matches full Premium without paying frontier-model prices at every step, saving ≈80%.

## The criteria, scored

| Configuration | Fidelity | Soundness | Completeness | Relevance | Calibration | Overall |
|---|--:|--:|--:|--:|--:|--:|
| Flagship (raw) | 5.4 | 8.0 | 5.6 | 8.1 | 7.6 | 5.4 |
| 9B, raw | 4.2 | 5.0 | 4.3 | 6.7 | 4.0 | 4.2 |
| 9B, harnessed | 5.5 | 5.3 | 5.5 | 6.3 | 5.9 | 5.4 |
| Optimum | 7.7 | 7.8 | 7.6 | 8.1 | 7.9 | 7.6 |
| Optimum + Premium | 8.4 | 8.4 | 8.6 | 8.5 | 8.7 | 8.4 |
| Premium | 8.6 | 8.7 | 8.7 | 8.6 | 9.0 | 8.6 |

*All figures 0–10. Fidelity measures whether the answer actually performed the assigned analysis. Overall is capped at Fidelity, so substance overrides polish.*

- **The harness adds execution, not polish** — raw models already score ~8 on Soundness and Relevance; they fall down on Fidelity and Completeness, and that is exactly what the harness raises.
- **The ranking is an execution ranking** — every configuration's Overall lands within a tenth of its own Fidelity score; the hard cap means execution, not fluency, was always the limiting factor.

**What each criterion measures**

- **Fidelity (0.35)** — Did it actually *execute* this technique's method and structure, or merely describe it, name-drop it, or write a generic essay? This is the test of whether the analysis was performed at all, which is why Fidelity sets the hard cap.
- **Soundness (0.25)** — Are the claims and inferences correct and logically valid, with evidence kept distinct from assumption? (Reasoning, not external fact-checking.)
- **Completeness (0.18)** — Does it surface what this technique should for *this* prompt, or leave major gaps?
- **Relevance (0.12)** — Does it engage *this specific* prompt and material, or drift into a generic treatise?
- **Calibration (0.10)** — Is it honest about uncertainty, explicit about assumptions, and free of overclaiming?

The weights descend in the order that determines whether an analysis is useful at all. Fidelity is heaviest because executing the method is the precondition for everything else; an answer that never runs the technique has nothing to be sound or complete *about*. Soundness is next — a faithful but invalid analysis actively misleads. Completeness sits in the middle: a sound-but-partial analysis is still useful. Relevance is lighter because being about the right thing is a low bar most answers clear, and genuine drift is already caught by Fidelity and Completeness. Calibration is lightest because it governs how an answer *presents* its limits rather than the substance of the analysis. The **hard cap** is the capstone of the same logic: the weighted average is ceilinged at the Fidelity score, so the other four axes can only pull a result down, never lift it above execution.

## What the numbers show

The first result is the private 9B comparison. A raw flagship model scores 5.4. The same task run through Ora's harness on a 9B model scores 5.4. That is not a claim that the 9B became the same model; it is a claim that the harness brought the working result up to raw-flagship quality. The practical consequence is the important part: a small private model can deliver frontier-grade analytical work on local hardware, with far lower cost and complete data privacy when run on-device.

The second result isolates what the harness does for a small model. The raw 9B scores 4.2; the harnessed 9B scores 5.4. That is a roughly 30% gain from the systems layer alone. The failure rate also drops from 41% to 8%, which means the harness is not merely polishing the prose. It is getting the model to actually perform the assigned operation more often.

The third result shows that the harness is not only rescuing weak models. The raw flagship scores 5.4, while the Premium harness reaches 8.6. That is roughly a 60% gain over the same basic intelligence tier when the model is put inside structured execution, adversarial review, correction, and consolidation. Ora magnifies frontier models too.

The fourth result is the most economically important one. Optimum uses inexpensive models in the harness and reaches 7.6, about 40% higher than the raw flagship's 5.4, with a 1% failure rate. The extra cost buys repeated passes, adversarial review, corrections, and reruns. That is the point of the architecture: cheap model calls, coordinated well, beat one expensive unstructured answer.

The fifth result shows where premium intelligence is most useful. Optimum + Premium keeps the low-cost Optimum pipeline but brings in the premium model at final consolidation. It reaches 8.4 at a 1% failure rate — among the lowest in the table. The pattern is like a senior reviewer examining the work product after the research staff has done the pipeline labor: the expensive intelligence is used where judgment is most leveraged.

The sixth result shows why Premium everywhere is wasteful. Full Premium reaches 8.6, only two-tenths above Optimum + Premium, while costing $228 per 100 queries instead of $40. In other words, using premium models throughout the pipeline spends roughly six times as much to gain almost nothing over using them at the key consolidation point.

The criteria explain why these comparisons line up the way they do. Raw models already score well on Soundness and Relevance: the raw flagship is around 8 on both. The bottleneck is execution. It scores 5.4 on Fidelity and 5.6 on Completeness because it often talks around the method instead of performing it. The harness lifts those execution axes: Optimum reaches 7.7 on Fidelity and 7.6 on Completeness, while Soundness and Relevance barely move.

That is why the ranking is effectively a Fidelity ranking. In every configuration, Overall lands within a tenth of Fidelity. The hard cap is doing exactly what it is supposed to do: a fluent answer cannot outrun the fact that it failed to execute the analysis. The order in the table is the execution order.

The reliability result follows from the same mechanism. Non-answers fall from 24% for the raw flagship and 41% for the raw 9B to 1–2% across the strongest harnessed configurations. Ora is not mainly turning good answers into slightly prettier answers. It is turning drift, gesture, and confabulation into completed analytical operations.

The trade is time. Raw models answer in about a minute. Harnessed runs take 17–24 minutes because they run multiple passes, reviews, corrections, and consolidation instead of one unreviewed generation. The table is measuring that exchange: minutes instead of seconds, in return for far higher quality and dramatically fewer non-answers.

## How we measured

Every answer was scored by a **two-judge panel** — a flagship-tier lead (Claude Opus) and a mid-tier second (Claude Sonnet). Each judge read the technique's specification and every configuration's answer — viewing the rendered diagram wherever one exists — to calibrate across the whole field, then scored its assigned lanes; the two judges' per-criterion scores are averaged.

The scale is **absolute, not a curve**: 9–10 is exemplary execution, ~5 is competent but flawed, and 2–3 means the answer only talks about the technique without performing it. The overall is the weighted average of the five criteria, then **capped at the Fidelity score** (the hard cap, 1.0) so a polished answer that never ran the method cannot outscore its own fidelity. A run is counted as a **failure** when its overall falls below 4.0 — the threshold below which an answer is closer to "talks about it" than to "competent."

The comparison covers **196 techniques** — analytical modes, interpretive lenses, and visual outputs — each answered from one shared prompt, so differences trace to the configuration and not to an easier question. Costs are priced at API-equivalent rates, so a model served under a subscription is never counted as free. The visual-output diagrams were regenerated through Ora's improved rendering and synthesis pipeline and re-scored by the same two-judge panel (2026-06-23); the analytical text was retained from the original run, so only the visual deliverable changed.

## What each configuration is

- **Flagship Model** — Claude Opus 4.8 called once: no harness, no tools, no pipeline. The "just use the best model" control.
- **Small Model, Raw** — Qwen 3.5 9B called once, no harness. The cheapest possible baseline and the control twin of *Small Local Only*.
- **Small Local Only** — Qwen 3.5 9B in every harness role: analyst, reviewer, utility, consolidation, and verification. Fully on-device.
- **Optimum** — Ora's low-cost mixed-model stack: Qwen 3.7 Plus and MiniMax M3 for large-model roles; Qwen 3.6 35B, Mimo v2.5, and Qwen 3.5 9B for fast/small/utility roles; Qwen 3.7 Plus for consolidation and verification.
- **Optimum + Premium** — the Optimum stack, but with Claude Opus 4.8 added for the final consolidation pass.
- **Premium** — Claude Opus 4.8 in the large, consolidation, and verification roles, with Claude Haiku 4.5 in the fast, small, and utility roles.

## Explore the capabilities

The comparison runs across the three kinds of work Ora produces:

## Reproducibility

Every run is captured — the prompt, the full output, token counts, cost, and wall-clock time — and every score is stored per technique and per configuration alongside the judge's notes for each criterion. To reproduce the capture set, download the [Ora Performance capture framework](/downloads/ora-performance-capture-framework.md), which includes the 196-entry prompt library and uses API calls only. To reproduce the ratings, run the separate [judging and scoring framework](/downloads/ora-performance-judging-framework.md) against those captures; it selects its two judge models from the Optimum configuration's large-model slots.

## Related

This paper measures the architecture argued for in [Harness Control Loop](/papers/harness-control-loop) and [Methodology](/papers/methodology); for how model choice is made inside a configuration, see [Model Selection](/papers/model-selection).
