What is a per-model grade?
A per-model grade scores how well one kit performs on one AI model, marked with whether it was measured from real runs or declared by the publisher.
Because a kit brings the instructions rather than the model, the same kit runs on Claude, ChatGPT, Gemini and Grok. That immediately raises the question a grade exists to answer: **does it run equally well on all of them?**
Usually not, and the honest thing is to say so per model rather than to claim uniform performance.
Measured or declared
Every grade carries its provenance, and this is the part that matters most:
| Measured | Declared | |
|---|---|---|
| Where it comes from | Real runs of this kit | The publisher's assessment |
| What it tells you | What happened | What the author expects |
| Weight it deserves | Evidence | A claim worth checking |
| Available when | The kit has run enough | Immediately, including at launch |
Both are useful and they are not equivalent. Publishing the distinction rather than one blended number is what makes an open publishing model workable: a new kit can say what its author believes without that being dressed up as data.
The practical rule: **a declared grade is a starting hypothesis, a measured grade is a result.** Weigh them differently before putting a worker on a job that reaches customers.
Why a kit performs differently across models
The differences that show up in unattended work are narrower than general benchmarks suggest, and mostly about consistency rather than raw capability:
- Instruction adherence over long context. Does the rule still hold on run four hundred, on an oddly shaped input?
- Restraint. Does it stop when information is missing, or produce a plausible answer to fill the gap?
- Format stability. A worker whose output another system reads needs the same shape every time.
- Synthesis across sources. Combining four apps is a different skill from extracting from one.
That last one is why grades diverge most on multi-app kits. Meeting Prep Assistant reaches calendar, email, meeting notes and CRM and has to produce one coherent document from all four. Engineering Pulse reads GitHub and summarises it. The first discriminates between models far more than the second, which is worth knowing when you read a grade that looks surprisingly flat.
How to use a grade
A grade narrows the field. It does not make the decision.
| Situation | What to do |
|---|---|
| Grades are close | Pick on cost or on the model you already use |
| One is clearly ahead, measured | Start there |
| One is clearly ahead, declared | Start there, then verify on your own data |
| The job touches customers | Run your own comparison regardless |
The verification step is cheap: send the same run to two models and read both outputs. Same instruction, same inputs, same day. Two runs of a four-app worker cost roughly 40 tool calls, about 8% of Free's daily 500.
You can also filter the directory by model to see only kits that fit the one you already use.
The failure mode: treating a grade as a guarantee
A grade is an aggregate over runs that were not your runs, on data that was not your data.
A kit graded highly for meeting prep was graded against whatever meetings the measurement covered. If your team writes notes in a house shorthand, or your CRM uses custom fields, the thing that will actually go wrong is specific to you and appears in no grade.
That is not a criticism of grades; it is the limit of what any aggregate can carry. Grades are for narrowing from four models to one or two. Your own comparison is for choosing between those.
The reasonable objection
"Publishers grade their own kits. Why would anyone declare a low grade?"
Fair, and it is precisely why the provenance mark exists rather than a single score. A declared grade is labelled as the author's claim, so a reader can discount it accordingly, and a measured grade sitting beside it is the check.
The incentive is also less perverse than it looks. A publisher who declares uniformly high grades gets installs followed by workers that underperform, receipts that show it, and a page carrying their name. The address is permanent, which cuts both ways: it accumulates reputation as well as history.
The honest caveat stands: on a kit with no measured grades yet, you are reading a claim, and the ten minutes spent on your own comparison is the answer.
When a grade should not drive the decision
- The job is trivial. Classifying into four buckets rarely discriminates between frontier models; pick on cost.
- Compliance already decided. If one provider is approved, a grade is information, not a choice.
- You have your own measurements. Yours beat aggregates, always.
- Output feeds another system. Format stability may matter more than the quality the grade reflects.
FAQ
What does a per-model grade on a kit page mean?
It scores how that kit performs on that specific model, and it is marked with where it came from: measured from real runs, or declared by the kit's publisher. It is meant to narrow your choice between models, not to replace testing on your own data.
What is the difference between a measured and a declared grade?
A measured grade comes from real runs of that kit and is evidence. A declared grade is the publisher's own assessment and is a claim worth checking. Both are shown, marked, so a reader can weigh them differently.
Why does the same kit perform differently on different models?
Mostly through consistency rather than raw capability: instruction adherence over long context, restraint when information is missing, format stability, and the ability to synthesise across several apps. Multi-app kits discriminate between models much more than single-source ones.
Should I trust a declared grade?
Treat it as a starting hypothesis rather than a result. It tells you what the author expects, which is genuinely useful for narrowing the field, and then a comparison on your own data is what settles it, especially if the worker reaches customers.
How do I test a model on my own job?
Send the same run to a different model and read both outputs, so the instruction, the inputs and the day are identical and the model is the only variable. It costs roughly one extra run, and model tokens bill at provider list price with no markup. See /pricing.