Ora’s harness provides the missing systems layer for AI: structured execution, adversarial review, error correction, drift control, and reruns around interchangeable models. It brings frontier-grade quality to small, locally run models, magnifies frontier-model abilities, and makes complex, multi-step cognitive automation reliable enough to fulfill the long-awaited promise of AI agents.

  • Capability: Ora makes interchangeable models execute structured analytical techniques and guided cognitive operations accurately, turning models from answer generators into working components inside a larger system.
  • Reliability: Ora reduces drift and prevents failure through adversarial review, fact checks, course corrections, and reruns when an operation falls below standard.
  • Automation: Ora makes multi-step AI work viable because its reliability architecture catches bad intermediate steps before errors compound.
  • Local power: Ora lets small private models reach raw-frontier performance on-device, reducing or eliminating cloud dependency.
  • Economics: Ora requires fewer user prompts to reach higher-quality answers while allowing more cost-effective models, reducing total costs.

The results

ConfigurationRating /10Failure rateCost / 100 queriesAvg. time / run (min)
Flagship Model — frontier model, raw5.423%$81.0
Small Model, Raw — 9B, no harness4.240%$0.061.2
Small Local Only — 9B in the harness5.58%$318.9
Optimum — Ora’s recommended setup7.62%$1322.3
Optimum + Premium8.32%$3923.4
Premium — maximum capability8.62%$22817.2

Ratings are 0–10. Fidelity measures whether the answer actually performed the assigned analysis. Overall is capped at Fidelity, so a fluent answer that does not execute the method cannot score high. Cost is the API-equivalent price per 100 queries; time is wall-clock minutes per run. The unrounded costs are $7.76, $0.06, $3.21, $13.49, $39.1494, and $227.7939 in table order. Four recovered subscription-inclusive cells use explicit API-equivalent estimates because their raw usage records did not survive.

  • Private 9B reaches flagship quality — a harnessed 9B slightly exceeds the raw flagship, making frontier-grade work possible on today’s 16GB local machines with complete data privacy.
  • The harness lifts small models — the same 9B improves ≈31%, from 4.2 raw to 5.5 harnessed.
  • The harness lifts frontier models too — the raw flagship rises ≈59%, from 5.4 raw to 8.6 inside the Premium harness.
  • Low-cost models beat raw flagship quality — Optimum improves on the raw flagship by ≈41%, with a 2% failure rate, using inexpensive models and adversarial review.
  • Premium works best as final review — Optimum + Premium gets most of the benefit by using frontier intelligence only at consolidation.
  • Premium everywhere wastes money — Optimum + Premium improves on the raw flagship by ≈54% and nearly matches full Premium while saving ≈83%.

The criteria, scored

ConfigurationFidelitySoundnessCompletenessRelevanceCalibrationOverall
Flagship (raw)5.48.05.78.17.65.4
9B, raw4.25.04.46.74.14.2
9B, harnessed5.65.35.66.45.95.5
Optimum7.77.87.78.18.07.6
Optimum + Premium8.48.48.68.68.78.3
Premium8.78.78.88.79.18.6

All figures 0–10. Fidelity measures whether the answer actually performed the assigned analysis. Overall is capped at Fidelity, so substance overrides polish.

  • The harness adds execution, not polish — raw models already score ~8 on Soundness and Relevance; they fall down on Fidelity and Completeness, and that is exactly what the harness raises.
  • The ranking is an execution ranking — every configuration’s Overall lands within a tenth of its own Fidelity score; the hard cap means execution, not fluency, was always the limiting factor.

What each criterion measures

  • Fidelity (0.35) — Did it actually execute this technique’s method and structure, or merely describe it, name-drop it, or write a generic essay? This is the test of whether the analysis was performed at all, which is why Fidelity sets the hard cap.
  • Soundness (0.25) — Are the claims and inferences correct and logically valid, with evidence kept distinct from assumption? (Reasoning, not external fact-checking.)
  • Completeness (0.18) — Does it surface what this technique should for this prompt, or leave major gaps?
  • Relevance (0.12) — Does it engage this specific prompt and material, or drift into a generic treatise?
  • Calibration (0.10) — Is it honest about uncertainty, explicit about assumptions, and free of overclaiming?

The weights descend in the order that determines whether an analysis is useful at all. Fidelity is heaviest because executing the method is the precondition for everything else; an answer that never runs the technique has nothing to be sound or complete about. Soundness is next — a faithful but invalid analysis actively misleads. Completeness sits in the middle: a sound-but-partial analysis is still useful. Relevance is lighter because being about the right thing is a low bar most answers clear, and genuine drift is already caught by Fidelity and Completeness. Calibration is lightest because it governs how an answer presents its limits rather than the substance of the analysis. The hard cap is the capstone of the same logic: the weighted average is ceilinged at the Fidelity score, so the other four axes can only pull a result down, never lift it above execution.

What the numbers show

The first result is the private 9B comparison. A raw flagship model scores 5.4. The same task run through Ora’s harness on a 9B model scores 5.5. That is not a claim that the 9B became the same model; it is a claim that the harness brought the working result to raw-flagship quality. The practical consequence is the important part: a small private model can deliver frontier-grade analytical work on local hardware, with far lower cost and complete data privacy when run on-device.

The second result isolates what the harness does for a small model. The raw 9B scores 4.2; the harnessed 9B scores 5.5. That is a roughly 31% gain from the systems layer alone. The failure rate also drops from 40% to 8%, which means the harness is not merely polishing the prose. It is getting the model to actually perform the assigned operation more often.

The third result shows that the harness is not only rescuing weak models. The raw flagship scores 5.4, while the Premium harness reaches 8.6. That is roughly a 59% gain over the same basic intelligence tier when the model is put inside structured execution, adversarial review, correction, and consolidation. Ora magnifies frontier models too.

The fourth result is the most economically important one. Optimum uses inexpensive models in the harness and reaches 7.6, about 41% higher than the raw flagship’s 5.4, with a 2% failure rate. The extra cost buys repeated passes, adversarial review, corrections, and reruns. That is the point of the architecture: cheap model calls, coordinated well, beat one expensive unstructured answer.

The fifth result shows where premium intelligence is most useful. Optimum + Premium keeps the low-cost Optimum pipeline but brings in the premium model at final consolidation. It reaches 8.3 at a 2% failure rate — about 54% above the raw flagship. The pattern is like a senior reviewer examining the work product after the research staff has done the pipeline labor: the expensive intelligence is used where judgment is most leveraged.

The sixth result shows why Premium everywhere is wasteful. Full Premium reaches 8.6, only three-tenths above Optimum + Premium, while costing about $228 per 100 queries instead of $39. In other words, using premium models throughout the pipeline spends nearly six times as much to gain little over using them at the key consolidation point.

The criteria explain why these comparisons line up the way they do. Raw models already score well on Soundness and Relevance: the raw flagship is around 8 on both. The bottleneck is execution. It scores 5.4 on Fidelity and 5.7 on Completeness because it often talks around the method instead of performing it. The harness lifts those execution axes: Optimum reaches 7.7 on both Fidelity and Completeness, while Soundness and Relevance barely move.

That is why the ranking is effectively a Fidelity ranking. In every configuration, Overall lands within a tenth of Fidelity. The hard cap is doing exactly what it is supposed to do: a fluent answer cannot outrun the fact that it failed to execute the analysis. The order in the table is the execution order.

The reliability result follows from the same mechanism. Non-answers fall from 23% for the raw flagship and 40% for the raw 9B to 2% across the strongest harnessed configurations. Ora is not mainly turning good answers into slightly prettier answers. It is turning drift, gesture, and confabulation into completed analytical operations.

The trade is time. Raw models answer in about a minute. Harnessed runs take 17–24 minutes because they run multiple passes, reviews, corrections, and consolidation instead of one unreviewed generation. The table is measuring that exchange: minutes instead of seconds, in return for far higher quality and dramatically fewer non-answers.

How we measured

The current scoring panel uses two independent direct-API judges from the Optimum configuration: Big 1 is qwen/qwen3.7-plus and Big 2 is minimax/minimax-m3. Each receives the same technique specification, prompt, answer, and rendered diagram where one exists; each scores independently, then their per-criterion scores are averaged.

The scale is absolute, not a curve: 9–10 is exemplary execution, ~5 is competent but flawed, and 2–3 means the answer only talks about the technique without performing it. The overall is the weighted average of the five criteria, then capped at the Fidelity score (the hard cap, 1.0) so a polished answer that never ran the method cannot outscore its own fidelity. A run is counted as a failure when its overall falls below 4.0 — the threshold below which an answer is closer to “talks about it” than to “competent.”

The comparison covers all 198 kind-qualified techniques — 60 analytical modes, 22 visual outputs, and 116 interpretive lenses — each answered from its one featured prompt across six configurations, for 1,188 accepted and scored cells. The two causal-dag identities and the two fishbone-diagram identities remain separate end to end. Costs are priced at API-equivalent rates, so a model served under a subscription is never counted as free. The finalized 27 affected cells retain current raw panel records; the other 1,161 legacy cells retain their existing two-judge aggregate ratings, and this paper does not claim that their raw judge payloads survive.

What each configuration is

  • Flagship Model — Claude Opus 4.8 called once: no harness, no tools, no pipeline. The “just use the best model” control.
  • Small Model, Raw — Qwen 3.5 9B called once, no harness. The cheapest possible baseline and the control twin of Small Local Only.
  • Small Local Only — Qwen 3.5 9B in every harness role: analyst, reviewer, utility, consolidation, and verification. Fully on-device.
  • Optimum — Ora’s low-cost mixed-model stack: Qwen 3.7 Plus and MiniMax M3 for large-model roles; Qwen 3.6 35B, Mimo v2.5, and Qwen 3.5 9B for fast/small/utility roles; Qwen 3.7 Plus for consolidation and verification.
  • Optimum + Premium — the Optimum stack, but with Claude Opus 4.8 added for the final consolidation pass.
  • Premium — Claude Opus 4.8 in the large, consolidation, and verification roles, with Claude Haiku 4.5 in the fast, small, and utility roles.

Explore the capabilities

The comparison runs across the three kinds of work Ora produces:

The full catalog of analytical modes, interpretive lenses, and visual outputs — every technique scored here — lives in the paper registry, grouped by territory and family.

Reproducibility

Every accepted cell retains its prompt attribution, full output, cost, and wall-clock time. The capture audit reports 1,188 valid and 0 affected: 4 verified, 20 attested, and 1,164 accepted unverified_legacy records. Historical execution health is narrower than row completeness: 792 accepted Ora traces are in scope, 0 retain step-health records, and all 792 are disclosed as historical missing health rather than inferred healthy runs. To reproduce the capture set, download the Ora Performance capture framework, which includes the full 198-entry prompt library and uses API calls only. To reproduce the ratings, run the separate judging and scoring framework against those captures with the two direct-API judges named above.

This paper measures the architecture argued for in Harness Control Loop and Methodology; for how model choice is made inside a configuration, see Model Selection.