Meta Audit Series

Yggdrasil III: Structural Provenance Is Not Trust

Gaia Research
August 29, 2026
Abstract
Yggdrasil III makes Trust Magnitude fairer across suites and unique skills. An emergency amendment recalibrates github-stars-own evidence to min(250, stars/250), without the former skill-count divisor, so every current 5-star skill remains above the 250 TM floor. Fusion-recipe rows still describe how a suite is assembled, but contribute 0 Trust Magnitude. Shared suite-repository evidence remains capped at 50 Trust Magnitude per component, and an S grade still needs a positive eligible benchmark result, verifier attestation, or peer review at the skill's own layer. These rules correct evidence accounting; they do not change stored stars or claim that a skill became less capable.

Trust Magnitude (TM) summarizes the positive evidence supporting a named skill. It is not a capability score, a performance ranking, or a verdict on usefulness. Yggdrasil III corrects how TM treats suites so that a suite's structure and shared repository standing cannot look like many independent confirmations.

The intent is balance, not punishment. Suites keep their composition, provenance, and eligible evidence. Unique skills and suite components are simply compared without repeatedly turning one shared signal into fresh corroboration. The rules changed TM calculations and public projections; they did not change stars.

Emergency amendment: repository-star recalibration

The initial Yggdrasil III projection exposed an overcorrection: github-stars-own divided repository adoption by up to four according to skillCountInRepo, leaving several established 5-star suites below the 250 TM floor. PR #1666 replaces that row formula with min(250, stars/250). Repository adoption now contributes directly, reaches its row cap at 62,500 stars, and no longer amplifies unrelated evidence. Same-source deduplication, the 50 TM suite-component baseline cap, and the independent S-witness safeguard remain unchanged.

A full gaia dev calibrate-trust-magnitude pass on the five current 5-star records produced:

SkillRecalibrated TMOverall Trust Grade
addy-osmani/agent-skills286.00A
garrytan/gstack331.59A
mattpocock/skills329.90S
obra/superpowers315.15A
ruvnet/ruflo290.00A

All five remain above the numeric 5-star floor. The four A outcomes are intentional: stars establish adoption magnitude but do not substitute for an eligible own-layer benchmark-result, verifier-attestation, or peer-review witness. This emergency calibration preserves current rank records while the broader per-type balance audit continues in issue #1665.

What was corrected

Previously, a fusion-recipe row could add a large number to TM. That row describes which skills make up a Fusion; it does not independently test or validate them. Yggdrasil III therefore keeps the row for provenance, graph traversal, and rank rules, but fixes its TM contribution at 0 and excludes it from scoring Evidence Type diversity.

QuestionBeforeYggdrasil III
What does a fusion recipe show?Structure and a numeric TM contributionStructure only
Does it add a scoring Evidence Type?It couldNo
Does the suite disappear from the graph?NoNo

Before and after evidence flow. Before Yggdrasil III, shared repository evidence and fusion structure could both increase component Trust Magnitude. After the change, shared suite evidence is capped at 50 TM per component, fusion structure contributes zero TM, and unique evidence remains eligible.

Figure 1. Structure remains visible, but only eligible evidence contributes to Trust Magnitude.

The suite-versus-unique balance, in plain language

Imagine a suite repository containing several components. When a component has repo-own or github-stars-own evidence pointing to that shared repository, the evidence can contribute a baseline. But the repository does not become new, independent proof each time the same source supports another component.

Yggdrasil III applies three simple rules:

This bounds repeated shared or structural evidence while preserving component-specific evidence. Unique skills receive no bonus, and suites receive no blanket penalty: eligible evidence follows the same scoring rules, with only a suite component's repeated shared-repository baseline capped.

Suite and unique-skill balance diagram. One suite repository provides each component with at most a 50 TM shared baseline. Component-specific evidence can add normally. Fusion-recipe links remain visible as structure but add zero TM. A unique skill is scored from its own eligible evidence.

Figure 2. Shared suite standing is a bounded baseline; specific evidence does the differentiating work.

The S-grade safeguard

An S grade now requires all three of the following:

1. TM of at least 250. 2. At least three distinct Evidence Types with positive scores. 3. At least one positive, eligible benchmark-result, verifier-attestation, or peer-review row recorded at the skill's own evidence layer.

Rejected benchmark rows, deranked verifier rows, zero-scoring rows, phantom rows, and inherited witnesses cannot satisfy the third requirement. Repository evidence may still contribute where its formula allows, but it cannot replace the eligible own-layer witness.

The safeguard asks whether strong TM has at least one direct observation of the skill. It does not say that a skill without such a witness is bad, broken, or incapable.

What changed in the public projection

On the reviewed 264-skill snapshot, the corrected calculation produced 0 S, 58 A, 82 B, 106 C, and 18 ungraded named skills. The graph, API, contributor pages, and named-skill pages were regenerated from that calculation.

Those grade movements describe the available corroborating evidence under the new rules. They do not describe a loss of capability. The Yggdrasil III implementation did not edit evidence rows or stored stars.

Follow-up: the resolver correction

Review then found a narrower resolver bug. The shared skill-map loader removed every frontmatter role field because it confused the RFC marker role: variant with an unrelated display-only role. As a result, variant components could be treated as graded origins during fusion-recipe origin resolution.

PR #1647 fixed the shared resolver and added regression coverage. After that fix, merged PR #1649 recalibrated the one affected stored TM projection, ruvnet/ruflo-v3:

CheckResult
Trust Magnitude216.00 → 186.00 (−30.00)
Overall Trust GradeA → A
Eight genuine role: variant entriesIndividual dry-runs were no-ops
Evidence rows, rank, and starsUnchanged

This follow-up synchronized the named-skill record and its generated API, search, and index projections. It was not new evidence, a demotion, or a star change.

What remains unchanged

Yggdrasil III does not remove suites, erase their provenance, or discount component-specific evidence. It does not change the Star Bar, suite membership, origin attribution, evidence records, rank, or stars. It changes how TM distinguishes structural/shared context from corroborating evidence.

The result preserves two ideas at once:

References
[2] Gaia Skill Tree. Trust Magnitude methodology.