Loading
Commits on Source 33
-
Coung Ngo authored
Safari doesn't focus <button> elements on click, and instead blurs whatever was previously focused. When a dropdown's search input was focused and the user clicked a button inside the dropdown (e.g. a listbox "Clear" reset button), that blur fired a 'focusin' targeting an element outside the bound container, so GlOutsideDirective closed the dropdown on mousedown, before the click could reach the button. Extend the existing mousedown-fallback logic (already used to handle click targets changing during text selection) to also cover 'focusin', so a focusin caused by the same mousedown that started inside the bound element isn't treated as an outside interaction. Co-Authored-By:Claude Sonnet 5 <noreply@anthropic.com>
-
Coung Ngo authored
Clearing lastMousedown on every focusin, not just click, meant any focus change during a press (clicking an unfocused search input, grabbing static text) wiped the stored target before the gesture finished. The later click outside then had nothing to fall back on and closed the dropdown, regressing text-selection-and-release-outside behavior that worked before the previous commit. Keep lastMousedown until click fires, since click happens after mouseup and the target must survive the whole gesture. Restrict the focusin fallback to while the mouse button is down, tracked by a new isMouseDown flag, so the original Safari mousedown-focus-loss case is still caught while unrelated focusin events (Tab, programmatic focus) are judged on their own target instead of a stale one. This also fixes right-click outside followed by Tab back into the dropdown, since a right-click fires no click event to clear the stored target. See discussion: !6339 (comment 3836317848) Co-Authored-By:
Claude Opus 5 (1M context) <noreply@anthropic.com>
-
Chad Lavimoniere authored
-
Chad Lavimoniere authored
Docs examples 10/14: Rework messaging component examples See merge request !6321 Merged-by:
Chad Lavimoniere <clavimoniere@gitlab.com> Approved-by:
Adam Ferch <aferch@gitlab.com> Approved-by:
Tim Noah <tnoah@gitlab.com> Reviewed-by:
GitLab Duo <gitlab-duo@gitlab.com> Co-authored-by:
Duo Developer <service_account_group_9970_1976e9ecde3c53c783b64edb9ae993eb@noreply.gitlab.com>
-
Adam Ferch authored
Every edit to the skill is currently unmeasured, including the seven commits in !6296 This is eleven prompts with pass criteria a human can grade in about ninety minutes, so a change to the skill or to the docs can be checked rather than argued about. Cases are derived from the skill's own instructions, and each one only earns its place if the skill changes the answer. That weights scope gating and page resolution over component choice, since any capable model already knows a modal blocks interaction. Ground truth for every check is verified against the repo and stated inline, so grading needs no Pajamas expertise and two reviewers should reach the same verdict. Co-Authored-By:
Claude Opus 5 (1M context) <noreply@anthropic.com>
-
Adam Ferch authored
A3's third check scored the same whether or not the gate fired: the fixture has no design system, so plain CSS is the natural answer and the point came for free. It also missed the failure it existed to catch, an agent that stops correctly and then writes `#dd2b0e` from memory, which is `--gl-color-red-500` but names no token and no component. The check now covers colour values, and A3's first check says what stopping means, since the skill's own stop text ends Pajamas rather than the task. B1 scored 3/3 with the skill disabled. `GlLoadingIcon` on the Spinner page is well enough known that the model answers from training, so the case measured the model. Swapped for `GlFilteredSearchSuggestionList` on `/components/filter`, where the slug and the component name share no word. D1's second check required every token to sit in the `-background-color-*` family, which failed a correct answer for naming the foreground and border tokens beside it. It's catching a wrong family, not extra members of the right one. Also softened the ablation rule to read per case. One case scoring flat means that case is measuring the model, not that the run is void: C1 and D1 both moved in the same pass B1 didn't. Co-Authored-By:Claude Opus 5 (1M context) <noreply@anthropic.com>
-
Adam Ferch authored
Co-Authored-By:Claude Opus 5 (1M context) <noreply@anthropic.com>
-
Adam Ferch authored
The old A3 put a GitHub remote in front of a check that stopped on any non-gitlab.com forge. That check is gone: the skill points at Pajamas rather than standing down, and the remote doesn't change the answer. Run against the current skill the case scored 0/3 for correct behavior, which makes it a trap rather than a test. The replacement guards the line that took its place. The skill says an installed component library is the signal and lists six by name, so the fixture installs `@acme/design-system`, which is none of them. If someone later rewrites that line as a closed enumeration, this case catches it. It's also the harder half of a pair. A2 has a `gitlab-org/` remote, so a model can reach the right answer by reasoning that a GitLab repo should mention Pajamas. A3 has no GitLab signal at all, so the only reason to raise Pajamas is the skill. Verified against the current skill: recognised the library, named Pajamas as the house system, offered to move the project, converted nothing, wrote no GitLab values, and applied the accessibility and content guidance that survives. 3/3. Also updates A2's and A4's ground truth, which described affiliation checks and a skill that stops. Neither exists now. Co-Authored-By:Claude Opus 5 (1M context) <noreply@anthropic.com>
-
Adam Ferch authored
Eleven cases against pajamas SKILL.md at cf4c5de3, Opus 5, scratch directories with the marketplace copy disabled. Overall 0.91. Bucket A 1.00, B 0.67, C 1.00, D 1.00, global checks clean on all eleven. Bucket A was 0.67 this morning and every skill edit behind that move was in the section bucket A tests. B, C and D didn't move across two rewrites of that section, which is the more useful signal. B3 is the only failure and it stays one until the docs agree with themselves. The ablation is recorded as not rerun rather than carried over, because it last ran against the previous skill and the previous B1, so the arm that says whether any of this measures the skill is currently unmeasured. Saying so beats a stale number. Two setup notes were wrong and the run is how I found out. The skill line count had drifted, so that's gone in favour of naming the two copies and saying to disable the marketplace one. And "grant WebFetch" was not enough: the scope question inspects dependencies and the A prompt asks for a build, so those cases need Write, Edit and Bash(git remote:*) too. Withholding Write doesn't make a case read-only, it makes the agent spend its turns narrating denied writes. Co-Authored-By:
Claude Opus 5 (1M context) <noreply@anthropic.com>
-
Adam Ferch authored
The old baseline had three problems. C1 was recorded 3/3 when it never stated Pajamas has no page for the component, which is its first check, so the overall was 0.91 rather than 0.88. It predated the B3 fix. And the ablation row said "not rerun", which left the only number that says whether any of this measures the skill permanently absent. Both arms now run at 7d3f7de7: skill 1.00, control 0.58. Twelve cases, $11, one commit, so the numbers are comparable to each other for the first time. The delta is the point and the headline hides it. Four scope cases are tautological, and the control scoring 0.08 there confirms the arm rather than measuring anything. Strip scope and it's 1.00 against 0.83, with the remaining signal in four places: the token family, the figure labels, the absence claim, and flagging a wrong path. Everything else the model does unaided with web access. That's worth stating plainly before anyone spends effort on the parts that don't move. Bucket E covers what the other four assume away. They all treat a page as prose, and 72 figures across the component pages aren't. The case asks about a Save/Cancel/Delete row, where "keep destructive buttons separate" exists only as `figure-img` labels inside `<do>` and `<dont>` and nowhere in the page's prose. Its second check asks for the quote rather than the rule, and that's deliberate. A model with no skill knows not to put Delete beside Cancel, says so confidently and cites the right page. What it doesn't do is reproduce a string that only exists in an element attribute. Measured twice per arm before settling on that wording. Also records two docs defects the run surfaced, neither of them skill problems: segmented-control missing from nav.json while button-group.md calls it deprecated, and the Filter page's when-to-use section being `<todo>Add when to use.</todo>`. Co-Authored-By:
Claude Opus 5 (1M context) <noreply@anthropic.com>
-
Adam Ferch authored
Four things were wrong. Eleven cases is now twelve. Ninety minutes was an estimate; both arms ran in about twenty and $11. "A and B are where failures are most likely" is backwards, since A is tautological and B is fixed. And the control arm is now part of a pass rather than an optional ablation, which is the biggest change to how this gets run and wasn't mentioned at all. The claim worth demoting is that the criteria carry their own ground truth so two people reach the same verdict. Three criteria have needed rewriting after a run showed two readings, and one case was graded wrong twice against the shipped wording. It's the right target and the doc should say when it isn't met. Also puts the six-of-twelve figure up front. It's the first thing that should shape how someone reads a result, and it was buried in the baseline. Co-Authored-By:Claude Opus 5 (1M context) <noreply@anthropic.com>
-
Both were left behind when bucket E and the measured control arm landed. The global-check count is out of twelve, matching the twelve cases and the Scoring section. B3 no longer fails: the skill reads the repo tree rather than nav.json, and the baseline has it at 3/3 in both arms. Raised in review on !6302
-
B1 measured 3/3 in both arms. The cause is architectural rather than a bad fixture: both arms get WebFetch and design.gitlab.com is public, so the control resolves any component-to-page mapping by search. No lookup case can discriminate on those terms. Its no-false-page check moves to B2, which the baseline scored on three checks, so the fourth is unmeasured. B3 stays as a guard on the button-group deprecation contradiction. The B1 row stays in the results as a record of the run and is excluded from the means. Eleven cases now, so the global-check denominator, the bucket B mean, the overall mean and the strip-scope figure all move.
-
The B1 row stays as a record of a run that happened, marked as cut and excluded from the means. Bucket B becomes 0.83, overall 0.55, and the strip-scope figure 0.81 across seven cases. The 0.58 that reproduced an independent control is kept as the twelve-case figure as run, so it is not confused with the recomputed eleven-case mean.
-
🤖 GitLab Bot 🤖 authored
-
Adam Ferch authored
Add a smoke test for the pajamas skill See merge request !6302 Merged-by:
Adam Ferch <aferch@gitlab.com> Approved-by:
Chad Lavimoniere <clavimoniere@gitlab.com> Reviewed-by:
GitLab Duo <gitlab-duo@gitlab.com>
-
Chad Lavimoniere authored
-
Chad Lavimoniere authored
Docs examples 11/14: Rework navigation component examples See merge request !6322 Merged-by:
Chad Lavimoniere <clavimoniere@gitlab.com> Approved-by:
Jeremy Elder <jelder@gitlab.com> Reviewed-by:
GitLab Duo <gitlab-duo@gitlab.com>
-
Jules Miocene authored
-
Chad Lavimoniere authored
Add decision-tree component with YAML-driven data See merge request !6248 Merged-by:
Chad Lavimoniere <clavimoniere@gitlab.com> Approved-by:
Chad Lavimoniere <clavimoniere@gitlab.com> Reviewed-by:
GitLab Duo <gitlab-duo@gitlab.com> Co-authored-by:
Duo Developer <service_account_group_9970_1976e9ecde3c53c783b64edb9ae993eb@noreply.gitlab.com> Co-authored-by:
Julia Miocene <jmiocene@gitlab.com>
-
🤖 GitLab Bot 🤖 authored
-
Chad Lavimoniere authored
-
Chad Lavimoniere authored
Docs examples 14/14: Rework progress and utility examples See merge request !6325 Merged-by:
Chad Lavimoniere <clavimoniere@gitlab.com> Approved-by:
Scott de Jonge <sdejonge@gitlab.com> Reviewed-by:
Scott de Jonge <sdejonge@gitlab.com> Reviewed-by:
GitLab Duo <gitlab-duo@gitlab.com> Co-authored-by:
Duo Developer <service_account_group_9970_1976e9ecde3c53c783b64edb9ae993eb@noreply.gitlab.com>
-
Vanessa Otto authored
Update dependency js-yaml to ^4.3.2 See merge request !6357 Merged-by:
Vanessa Otto <votto@gitlab.com> Approved-by:
Vanessa Otto <votto@gitlab.com> Co-authored-by:
GitLab Renovate Bot <gitlab-bot@gitlab.com>
-
Chad Lavimoniere authored
-
Chad Lavimoniere authored
Match product code styling in live-example previews See merge request !6352 Merged-by:
Chad Lavimoniere <clavimoniere@gitlab.com> Approved-by:
Sascha Eggenberger <seggenberger@gitlab.com> Reviewed-by:
GitLab Duo <gitlab-duo@gitlab.com>
-
Vanessa Otto authored
Update dependency @clack/prompts to ^1.8.1 See merge request !6356 Merged-by:
Vanessa Otto <votto@gitlab.com> Approved-by:
Vanessa Otto <votto@gitlab.com> Co-authored-by:
GitLab Renovate Bot <gitlab-bot@gitlab.com>
-
Chad Lavimoniere authored
-
Chad Lavimoniere authored
Docs examples 13/14: Rework avatar, badge, label, and token examples See merge request !6324 Merged-by:
Chad Lavimoniere <clavimoniere@gitlab.com> Approved-by:
Adam Ferch <aferch@gitlab.com> Reviewed-by:
GitLab Duo <gitlab-duo@gitlab.com> Co-authored-by:
Duo Developer <service_account_group_9970_1976e9ecde3c53c783b64edb9ae993eb@noreply.gitlab.com> Co-authored-by:
GitLab Bot <gitlab-bot@gitlab.com>
-
Chad Lavimoniere authored
-
Chad Lavimoniere authored
Docs examples 12/14: Rework table, card, and dashboard panel examples See merge request !6323 Merged-by:
Chad Lavimoniere <clavimoniere@gitlab.com> Approved-by:
Scott de Jonge <sdejonge@gitlab.com> Approved-by:
Ezekiel Kigbo <3397881-ekigbo@users.noreply.gitlab.com> Reviewed-by:
Scott de Jonge <sdejonge@gitlab.com> Reviewed-by:
GitLab Duo <gitlab-duo@gitlab.com>
-
Paul Gascou-Vaillancourt authored
Fix GlOutsideDirective closing dropdowns on mousedown in Safari See merge request !6339 Merged-by:
Paul Gascou-Vaillancourt <pgascouvaillancourt@gitlab.com> Approved-by:
Paul Gascou-Vaillancourt <pgascouvaillancourt@gitlab.com> Approved-by:
Thomas Hutterer <thutterer@gitlab.com> Reviewed-by:
Paul Gascou-Vaillancourt <pgascouvaillancourt@gitlab.com> Co-authored-by:
Coung Ngo <cngo@gitlab.com>
-
🤖 GitLab Bot 🤖 authored