Add a smoke test for the pajamas skill
The cases are worth arguing with, and the open questions at the bottom are still live, but none of them need settling before this lands.
Eleven prompts with pass criteria, run twice — once with the skill and once without — and graded by a human. About twenty minutes and $11 a pass. It answers something we currently guess at: when we edit the skill or the docs, did it help?
Why these cases and not others
A case only earns its place if the skill changes the answer. Any capable model knows a modal blocks interaction, so that case measures the model rather than our documentation. Page lookups turned out to be the same: with web access the control resolves any component-to-page mapping by searching the site, however obscure the slug. What it doesn't do unaided is doubt a path someone cited, read a rule that lives only in a figure label, or pick the confirm primary background out of the right token family and notice it's a neutral rather than a blue.
So this weights scope gating and the behavior around a lookup over the lookup itself, which is the opposite of where the intuition points.
The cases, and where they stand
Each scope case runs the same prompt, "Build a settings form with a name field and a save button", in a different fixture.
| What it tests | Skill | Control |
|---|---|---|
| Scope | 1.00 | 0.08 |
| No dependencies, no remote: apply Pajamas | 3/3 | 0/3 |
@mui/material, gitlab-org/ remote: name the gap, offer the move, convert nothing |
3/3 | 0/3 |
| A component library the skill doesn't name, no remote: same answer | 3/3 | 0/3 |
--brand-* properties only: apply Pajamas, a prefix isn't a design system |
3/3 | 1/3 |
| Page resolution | 1.00 | 0.83 |
An issue citing /components/form/select, which doesn't exist |
3/3 | 2/3 |
| Where the segmented control is documented | 3/3 | 3/3 |
| Grounding and refusal | 1.00 | 0.83 |
Reviewing GlClipboardButton, exported with no guidance page |
3/3 | 2/3 |
Reviewing GlNotificationBar, which doesn't exist |
3/3 | 3/3 |
| Tokens | 1.00 | 0.83 |
| The background token for a primary confirm button | 3/3 | 2/3 |
A .status-pill with a hex, a px size, and a shadow standing in for a border |
3/3 (+1) | 3/3 |
| Reading the whole page | 1.00 | 0.67 |
| A Save/Cancel/Delete row, where the rule lives only in figure labels | 3/3 | 2/3 |
| Overall | 1.00 | 0.55 |
Read the delta, not the score. The four scope cases are tautological — that behavior exists only because the skill defines it, so the control confirms the arm rather than measuring anything. Strip scope out and it's 1.00 against 0.81, with the remaining signal in four places: the token family, the figure labels, the absence claim, and flagging a wrong path. Three of the eleven score identically in both arms.
What the runs found
The suite caught a defect in !6296 (merged) before that MR merged. The skill told agents nav.json "lists every guidance page" and to resolve paths from it. It's a hand-maintained sidebar, segmented-control is missing from it while being live on the site and in the repo, and an agent followed the instruction and reported the page as undocumented. 0/3 with the skill, 3/3 without. Fixed by reading the repo tree instead.
A twelfth case was cut after the run. It asked which page documents GlFilteredSearchSuggestionList and scored 3/3 in both arms, and the cause is architectural rather than a bad fixture: both arms are granted WebFetch and the site is public, so the control resolves any mapping by searching it. Its no-false-page check moved to the /components/form/select case.
Two docs defects are still open, neither of them skill problems. segmented-control.md carries live guidance while button-group.md:123 calls the component deprecated. And the Filter page's when-to-use section is literally <todo>Add when to use.</todo>.
What would help
- Three cases don't discriminate. Under this suite's own rule they should be cut, but a case that catches a regression earns its place even if it never moves. I don't think that's settled.
- Whether the criteria are unambiguous enough that two of us grade identically. Three weren't, and one was graded wrong twice.
- The rerun rule is asymmetric: only failures get rerun, so a flaky pass stays a pass. Three runs per case is the shape I'd want.
- A skill arm at 1.00 means the suite can't tell a good skill from a better one. Worth deciding whether that's a problem to fix or a place to stop.
Related to #3598.