What is a per-model grade?

A per-model grade scores how well one kit performs on one AI model, marked with whether it was measured from real runs or declared by the publisher.

Because a kit brings the instructions rather than the model, the same kit runs on Claude, ChatGPT, Gemini and Grok. That immediately raises the question a grade exists to answer: **does it run equally well on all of them?**

Usually not, and the honest thing is to say so per model rather than to claim uniform performance.

Measured or declared

Every grade carries its provenance, and this is the part that matters most:

MeasuredDeclared
Where it comes fromReal runs of this kitThe publisher's assessment
What it tells youWhat happenedWhat the author expects
Weight it deservesEvidenceA claim worth checking
Available whenThe kit has run enoughImmediately, including at launch

Both are useful and they are not equivalent. Publishing the distinction rather than one blended number is what makes an open publishing model workable: a new kit can say what its author believes without that being dressed up as data.

The practical rule: **a declared grade is a starting hypothesis, a measured grade is a result.** Weigh them differently before putting a worker on a job that reaches customers.

Why a kit performs differently across models

The differences that show up in unattended work are narrower than general benchmarks suggest, and mostly about consistency rather than raw capability:

That last one is why grades diverge most on multi-app kits. Meeting Prep Assistant reaches calendar, email, meeting notes and CRM and has to produce one coherent document from all four. Engineering Pulse reads GitHub and summarises it. The first discriminates between models far more than the second, which is worth knowing when you read a grade that looks surprisingly flat.

How to use a grade

A grade narrows the field. It does not make the decision.

SituationWhat to do
Grades are closePick on cost or on the model you already use
One is clearly ahead, measuredStart there
One is clearly ahead, declaredStart there, then verify on your own data
The job touches customersRun your own comparison regardless

The verification step is cheap: send the same run to two models and read both outputs. Same instruction, same inputs, same day. Two runs of a four-app worker cost roughly 40 tool calls, about 8% of Free's daily 500.

You can also filter the directory by model to see only kits that fit the one you already use.

The failure mode: treating a grade as a guarantee

A grade is an aggregate over runs that were not your runs, on data that was not your data.

A kit graded highly for meeting prep was graded against whatever meetings the measurement covered. If your team writes notes in a house shorthand, or your CRM uses custom fields, the thing that will actually go wrong is specific to you and appears in no grade.

That is not a criticism of grades; it is the limit of what any aggregate can carry. Grades are for narrowing from four models to one or two. Your own comparison is for choosing between those.

The reasonable objection

"Publishers grade their own kits. Why would anyone declare a low grade?"

Fair, and it is precisely why the provenance mark exists rather than a single score. A declared grade is labelled as the author's claim, so a reader can discount it accordingly, and a measured grade sitting beside it is the check.

The incentive is also less perverse than it looks. A publisher who declares uniformly high grades gets installs followed by workers that underperform, receipts that show it, and a page carrying their name. The address is permanent, which cuts both ways: it accumulates reputation as well as history.

The honest caveat stands: on a kit with no measured grades yet, you are reading a claim, and the ten minutes spent on your own comparison is the answer.

When a grade should not drive the decision

FAQ

What does a per-model grade on a kit page mean?

It scores how that kit performs on that specific model, and it is marked with where it came from: measured from real runs, or declared by the kit's publisher. It is meant to narrow your choice between models, not to replace testing on your own data.

What is the difference between a measured and a declared grade?

A measured grade comes from real runs of that kit and is evidence. A declared grade is the publisher's own assessment and is a claim worth checking. Both are shown, marked, so a reader can weigh them differently.

Why does the same kit perform differently on different models?

Mostly through consistency rather than raw capability: instruction adherence over long context, restraint when information is missing, format stability, and the ability to synthesise across several apps. Multi-app kits discriminate between models much more than single-source ones.

Should I trust a declared grade?

Treat it as a starting hypothesis rather than a result. It tells you what the author expects, which is genuinely useful for narrowing the field, and then a comparison on your own data is what settles it, especially if the worker reaches customers.

How do I test a model on my own job?

Send the same run to a different model and read both outputs, so the instruction, the inputs and the day are identical and the model is the only variable. It costs roughly one extra run, and model tokens bill at provider list price with no markup. See /pricing.