
Executes the full automated test suite, collects pass/fail counts and coverage deltas, and surfaces any newly introduced failures with concise root-cause notes.

Skill rank progression over time. Hover for details.
All named implementations attributed to @garrytan in the registry.

Executes the full automated test suite, collects pass/fail counts and coverage deltas, and surfaces any newly introduced failures with concise root-cause notes.

Builds a complete design system by researching the product and competitors, then proposing a coherent package of typography, colours, spacing, layout, and motion — persisted as a DESIGN.md file that serves as the project's authoritative design source of truth.

Drives a browser to a target URL, navigates multi-step user journeys, and captures screenshots or structured observations for downstream verification or reporting.

Launches a new headless or headed browser session with the Gstack environment variables and extension profile pre-loaded, ready for automation commands.

Injects pre-authenticated session cookies into the browser context so subsequent automation steps can access gated pages without a manual login flow.

Automates the final production shipping stages — merging a PR, monitoring CI/deploy completion, and verifying live site health through canary checks — picking up where /ship leaves off with safety gates at each step to prevent broken deployments reaching users.

Provisions the deployment environment by creating secrets, configuring environment variables, and running infrastructure-as-code init steps before the first deploy.

Comprehensive, interactive architecture and implementation review before coding begins, systematically evaluating scope, architecture, code quality, test coverage, and performance through structured questioning — synthesising findings into actionable tasks and an explicit not-in-scope section.

Pre-landing code review combining structured checklist analysis with specialist subagents covering testing, security, and performance — plus adversarial review from both Claude and Codex — to catch SQL safety issues, LLM trust boundary violations, conditional side effects, and structural problems before merging.

Activates a conservative execution profile that pauses before irreversible actions, requests explicit confirmation for destructive operations, and logs all side effects.

Sets a change-freeze flag that blocks non-critical commits and PR merges until explicitly lifted, protecting release branches or post-incident windows.

Applies configurable content and output guardrails to agent responses, flagging or blocking unsafe outputs and logging violations with structured evidence for audit.

Clears the active change-freeze flag and restores normal merge permissions, logging the unfreeze event with a timestamp and justification.

Synthesises commit history, PR comments, and issue notes into a written sprint retrospective covering wins, misses, root causes, and action items.

The definitive autonomous "Founder mode" review and decision suite. An auto-review pipeline that reads the full CEO, design, engineering, and DX review skills from disk and runs them sequentially with auto-decisions using 6 decision principles.

Infrastructure-first security audit focusing on secrets archaeology, dependency supply chain, and CI/CD security. Includes OWASP Top 10, STRIDE threat modeling, and active verification with daily (zero-noise) and monthly (comprehensive) scan modes.

Automated end-to-end deployment workflow that merges the base branch, runs tests, reviews the diff, bumps the VERSION file, updates the CHANGELOG, commits, pushes to the remote, and creates a pull request in a single command.

Verify a research claim or academic citation by tracing it through publication → methodology → raw data → independent replication. Routes through perplexity-research for the actual web lookup, then formats results as a citation-checked brain page. Use when a book/article/conversation cites a study and you want to confirm the claim is real, replicated, and accurately characterized.

Systematic claim-by-claim verification for any content before it ships. Modeled on professional fact-checking desks (The New Yorker, ProPublica, IFCN standards): extract every verifiable claim, check each against live citable sources (never training data), assign a 6-level confidence status, apply corrections, and produce a scored pass/fail report. Includes a data-derived-claims gate for outputs produced FROM the brain or a database: PRODUCER ≠ VERIFIER (re-derive each claim via a different query path) and AFFILIATION ≠ AUTHORSHIP (person→thing claims resolve through typed edges), with delivery hard-blocked on unsupported claims.

Generates structured documentation using the Diataxis framework — tutorials, how-to guides, reference materials, and explanations — by thoroughly researching the codebase before writing, tailored to the needs of different reader types.

Reads the diff between two version tags and drafts concise, user-facing release notes following the Gstack changelog format.

Converts markdown or structured data into a polished PDF document with consistent heading styles, table formatting, and page layout ready for distribution.

Fetches target URLs with a headless browser, parses structured data from rendered HTML, and returns clean JSON or markdown ready for downstream analysis or ingestion.

Universal archivist for personal file archives (Dropbox/B2/Gmail-takeout/local-mount/hard-drive-dump). Filters for high-value content (the user's own writing, ideas, relationships) and surfaces it interactively. REFUSES TO RUN without an explicit gbrain.yml `archive-crawler.scan_paths:` allow-list.

Feed and whole-publication ingestion: turn an entire blog, newsletter, or RSS/Atom archive into brain source pages. Covers feed discovery, pagination walking, normalization to a common article shape, canonical-URL dedup, idempotent re-runs, 429 pacing, and empty-husk repair. This is the PUBLICATION-scope skill — a single article URL routes to idea-ingest instead. Per-article enrichment hands off to the brain-ingest-gate skill; public posts only (gated content is skipped, never worked around).

End-to-end discipline for turning any large data source (audio libraries, email takeouts, document corpora, chat exports, API dumps) into brain pages at scale. The lifecycle spine: SCHEMA → ACCESS → TRIAL → EVALUATE → IMPROVE → CODIFY → TEST → SKILLIFY → BULK → MONITOR. State is tracked in a durable JSON manifest (see MANIFEST-PATTERN.md) so any crash, session boundary, or subagent fan-out resumes from ground truth instead of memory.

Tiered LLM extraction pattern for large corpus processing (email archives, document dumps, transcript libraries). A utility-tier model triages and classifies at speed; the reasoning tier does the default deep read; the deep tier is the escalation for the highest-value content. Prevents spending deep-tier money on noise while ensuring the important content gets the best eyes. A deterministic privacy wall runs before any LLM call.

Runs a standardised prompt suite across multiple model versions, records latency and quality scores, and produces a ranked comparison table to guide model selection.

Author an eval for an existing skill from its REAL usage history, not its spec. Mine invocations from the brain's conversation archive (conversations/) and per-harness session transcripts — a user correction after an invocation is the gold signal — then synthesize an eval_contract plus 4-8 replayable cases with honesty labels (SPEC-DERIVED vs HISTORY-IMPLIED) and stage the result at skills/<name>/eval/autobench-<date>.md as PENDING-HUMAN-APPROVAL. Never rewrites SKILL.md. Ships two guard companions: panel integrity (multi-model judging must prove each provider actually responded) and the fail-improve taxonomy (logged LLM-fallback cases convert to deterministic code over time).

Web performance benchmarking that captures baseline metrics, compares current performance against those baselines, and identifies regressions in load times, Core Web Vitals, and bundle sizes across specified pages.

Rigorous product strategy and scope review in four modes — SCOPE EXPANSION, SELECTIVE EXPANSION, HOLD SCOPE, and SCOPE REDUCTION — evaluating architecture, security, data flows, testing, and edge cases to ensure plans ship at the highest standard before implementation begins.

Take any book (EPUB/PDF), produce a personalized chapter-by-chapter analysis. Each chapter is preserved in detail (The Chapter) and mirrored back to the reader's actual life (The Mirror) using brain context. The mirror observes and resonates — a friend pointing out parallels, NOT a consultant rearranging the reader's life, NOT a therapist assigning homework. The reader decides what to do about it. Layout is a top-aligned HTML table or stacked sections, never a bare markdown pipe table (pipe tables center-misalign uneven columns). Output is a single brain page at media/books/<slug>-personalized.md plus an optional PDF via brain-pdf.

Read a book, article, transcript, or case study through the lens of a specific strategic problem you're facing. Produces an applied playbook that maps the source onto the problem and gives short/medium/long-term recommendations. NOT for general book summaries.

Performs the brain-first read, enrich, write, attribution, and backlink cycle for a persistent knowledge base.

Bootstraps the GBrain knowledge store for a new project: creates the index structure, ingests seed documents, and validates retrieval with a sample query.

Incrementally syncs new documents and updated pages into the GBrain index, deduplicates embeddings, and reports ingestion counts and any errors.

Post-deployment monitoring that captures pre-release baseline screenshots, then continuously watches pages for console errors, performance regressions, and broken links — designed to surface failures within the first 10 minutes so problems are caught before they reach users at scale.

Captures thoughts and content into a persistent personal knowledge brain through one idempotent ingestion entrypoint.

Import AI-assistant chat exports (ChatGPT, Claude, Perplexity) and agent session transcripts into the brain as one dated page per conversation under conversations/, validate each page against the native conversation parser, extract facts via the native conversation-facts flow, and keep the archive gap-free with a detect-and-backfill loop. Then answer archive questions: "when did I first discuss X", trace how an idea evolved across past conversations, pull a specific thread.

Ingest meeting transcripts from ANY meeting recorder into brain pages with attendee enrichment, entity propagation, and timeline merge. One unified pipeline: normalize the source into a standard transcript record, split multi-meeting recordings, resolve speakers by evidence, create the page, pass every surprising claim through the consistency check (transcript + brain + plausibility), enrich every entity, then run the verification checklist — substance AND sequence. A meeting is NOT fully ingested until the enrich skill has processed every entity AND the verification checklist passes, including the sequence verify (PASS or explicit user waive).

Opt-in ambient signal capture. After explicit enablement, applies on substantive inbound messages to detect original thinking and entity mentions. Use an authorized sub-agent where supported; otherwise detect inline. Aim to never block the main response.

Build a TYPED citation/reference graph over an ingested corpus — not just embeddings. Flat similarity retrieval cannot tell you that document A *overrules* B, *distinguishes* C, or *relies_on* D. This skill extracts every inter-document reference, classifies the edge TYPE with LLM judgment, and writes first-class typed edges via `gbrain link`, so `gbrain graph-query --type` can walk the argument ("everything this brief relies on, minus anything overruled since"). Every cite-heavy corpus is the same shape: law, academic papers, patents, regulatory filings, a book's bibliography.

Spins up a structured Claude-versus-Codex debate over a proposed implementation, with each agent mounting adversarial critiques, to surface hidden design flaws before code lands.

Synthesizes and curates raw knowledge concepts into a grounded, tiered intellectual map.

Trace one idea's evolution through the brain: first mention, best articulation, related concepts, reversals, contradictions, abandoned branches, and the current live version. Use for single-idea conceptual lineage, not broad concept-map synthesis or structured entity metrics.

Token-hygiene audit of the always-loaded context stack — CLAUDE.md, AGENTS.md, auto-memory MEMORY.md, and the bootstrap-rendered identity files (SOUL.md, USER.md, ACCESS_POLICY.md, HEARTBEAT.md) or their harness equivalents. Finds redundancy, contradictions, stale content, compression candidates, and skill-extraction candidates; produces a ranked action list sorted by token savings with a risk class per finding. REPORT-ONLY: this skill never edits any audited file. Recommendations for bootstrap-rendered files target the interview answer bank / templates, never the rendered output. Judging routes through `gbrain eval cross-modal` (single cheap model by default; full multi-model panel is explicit opt-in).

Reads a saved context snapshot and reconstructs a warm session state, surfacing the last decision point and pending tasks so work can resume without recap.

Compresses the current session context into a compact summary file that can be restored later, enabling long-running workflows to survive context-window limits.

Structured data research: search sources, extract structured data, archive raw sources, maintain canonical tracker pages, deduplicate. Parameterized via YAML recipes for investor updates, donations, company updates, or any email-to-structured-data pipeline.

Deep-research a topic end to end and produce a permanent, reusable knowledge asset: archive every primary source verbatim (gated by the user's privacy/retention posture), write one 1:1 summary per source, then synthesize a single self-contained compendium page. Depth is a dial (base synthesis → grounded primaries → books + counter-canon → saturation), each level an idempotent superset of the one below. Distinct from data-research (structured trackers) and perplexity-research (web deltas): this produces prose knowledge synthesis backed by an archived source corpus.

Generates production-quality, Pretext-native HTML/CSS with real text reflow and dynamic layout from approved design mockups, CEO plans, or user descriptions — producing fully self-contained files with computed heights, responsive behaviour, and zero external dependencies.

Runs a structured UX audit over a product interface, scoring layout clarity, affordance, and accessibility against the Gstack design rubric to surface actionable improvements.

Rapid design exploration that generates multiple AI design variants, opens a comparison board for the user, collects structured feedback, and iterates until a preferred visual direction is reached.

Interactive, designer-led audit of UI/UX plans rating seven dimensions — information architecture, interaction states, user journey, AI slop risk, design system alignment, responsive/accessibility, and unresolved decisions — before implementation begins, with optional visual mockup generation.

Interactive multi-pass DX review for developer-facing products — APIs, CLIs, SDKs, and libraries — scoring eight UX dimensions from onboarding to error messaging via persona discovery and competitive benchmarking, producing an actionable improvement plan rather than just a score.

Audits CLI ergonomics, API surface clarity, and onboarding friction from a developer perspective, producing a scored report with prioritised fixes.

Ghostwrite content in a specific person's voice from a VALIDATED voice profile — tweets, replies, short posts, launch copy, recruiting blurbs, emails. Loads the subject's voice profile (people/<slug>-voice) plus first-party context from the brain, drafts 2-3 options in-register, then runs a hard voice-fidelity self-check before showing anything. Includes the profile BUILDER: if no validated profile exists, drafting hard-stops and this skill walks the corpus-to-fingerprint build instead. Never auto-posts.

Compress an agent's routing file (RESOLVER.md or AGENTS.md) by converting granular skill-per-row tables into functional-area dispatchers. Each area lists sub-skills in a "(dispatcher for: ...)" clause. The LLM reads one area entry and routes to the correct sub-skill. Proven via held-out A/B eval: dispatcher pattern outperforms naive pipe-table compression.

The installable GBrain router and capstone for knowledge capture, operations, and concept synthesis.

Scans the workspace for outdated dependencies, runs upgrades within semver-compatible bounds, re-runs tests, and commits a clean dependency bump PR.

The legendary fusion of Garry Tan's complete agent discipline library — browser QA, security auditing, design exploration, vertical-slice planning, and founder-mode orchestration — unified into a single autonomous product-development workflow.

Systematic root-cause debugging enforcing an Iron Law — no fix without first identifying root cause — guiding through four phases: investigation, analysis, hypothesis formation, and verified implementation with evidence gathering and pattern matching before any code changes.

Before fixing a slow/stale/timeout alert, measure the step yourself. Kill the theory with a stopwatch, not a code change. Measure-first ops triage for temporal alerts (stale, timeout, freshness, wedged, N hours behind) from gbrain doctor, autopilot, sync, and cron monitors — runs BEFORE any timeout raise, threshold change, or pipeline rewrite.

Generates a one-page project landing report from open issues, recent commits, and milestone progress, giving stakeholders a quick read on health and next steps.

Reads the active memory store, consolidates new observations from the current session, deduplicates stale entries, and writes back an updated, ranked knowledge base for future sessions.

Unified Minions skill for both deterministic shell jobs and LLM subagent orchestration. Replaces the older `gbrain-jobs` routing intent. Use when: submitting gbrain jobs, shell/background tasks, spawning subagents, checking progress, steering running work, pausing/resuming, parallel fan-out. One durable, observable, steerable queue interface. Also carries the durable-execution doctrine for any operation expected to exceed ~2 minutes: capability ladder, deadman checks that verify the result was reported, and content-addressed stage checkpoints for expensive pipelines.

YC-style startup and builder brainstorming. Startup mode uses six forcing questions (demand reality, status quo, desperate specificity, narrowest wedge, observation, future-fit) to expose reality. Builder mode focuses on design thinking for side projects and hackathons.

Wires a new MCP server into the Gstack agent environment, validates the tool manifest, and demonstrates round-trip invocation through a test prompt.

Takes a draft plan or system prompt, identifies vague or ambiguous instructions, and rewrites them to reduce hallucination and improve task completion rate.

Self-evolving skill optimization via SkillOpt-paper-grounded text-space optimizer.

Browser-driven web application testing that explores pages as a real user, documents bugs with annotated screenshots, fixes issues with atomic commits and re-verification, and produces structured reports with before/after evidence and health scores showing quality improvement.

Runs the scoped end-to-end test suite for a single feature or route without launching the full QA pipeline, for fast targeted regression checks.

Publish the user's own gbrain over MCP so other devices, desktop apps and cloud agents can reach it. `gbrain mcp expose` installs and signs in Tailscale, publishes the running `gbrain serve --http` on the tailnet (HTTPS, tailnet-only by default; Funnel only when a client lives in a vendor cloud), keeps the server alive as a user service, and hands back the MCP URL. Then grant one least-privilege client per consumer, install the handoff inside that client, and verify a real memory round trip.

Converts a freeform prompt, repo pattern, or workflow description into a complete, registry-ready named skill: writes the SKILL.md definition, populates frontmatter fields, and opens a PR for review.
Want to add more skills?
Register your repo →