Where models contribute
Four model providers contribute to grading. Their outputs are combined around a versioned rubric covering relevance, reasoning, clarity, accuracy and concision; no single model controls the verdict.
Experiment 02 / Logosmose / 2026
A written-reasoning experiment built to make the question, constraints, rubric and result inspectable rather than asking users to trust a single opaque judge.
01 / Hypothesis
Personal experience with inconsistent human evaluation made the trust problem concrete. The defensible product claim became narrower: evaluate a written response against a public rubric, not judge a person’s intelligence.
02 / Smallest loop
Reveal one difficult question.
Write for five minutes, within 2,000 characters.
Evaluate the response across five public criteria.
Return the score, criterion breakdown and rank.
One response. Four model opinions. One inspectable result.
Core mechanismProvider diversity reduces dependence on one opaque judge.
Control boundaryThe system scores one response against a rubric, not a person’s intelligence.
03 / Division of labour
Four model providers contribute to grading. Their outputs are combined around a versioned rubric covering relevance, reasoning, clarity, accuracy and concision; no single model controls the verdict.
I defined the product claim, rubric, constraints, aggregation rules and acceptable language. Users decide whether the result is useful; the system does not claim to measure intelligence.
04 / Product evidence
These production screens show the core loop now available for free participation and training. They demonstrate interface behavior, not user traction or grading validity.


Reality check / the business model changed
The original financial dimension created legal exposure, gambling associations and payment-processor friction. I removed it. That decision left a smaller free loop capable of testing trust and repeat participation before adding commercial complexity.
Whether the free core loop creates voluntary repeat use.