Skip to main content
Stenion
The rulebook

Methodology

Every formula, threshold, and weight Stenion uses — extracted directly from the shipped adapter code. It exists so anyone, including the protocols being scored, can verify and challenge the rules, not just the output. Code and this document are not allowed to drift.

View source on GitHub

Stenion Scoring Methodology

This document is the source of truth for how every safety factor is calculated. It exists so that anyone — including the protocols being scored — can see, verify, and challenge the actual rules, not just the output numbers. Every formula below is extracted directly from the shipped code (currently adapters/blend/ and adapters/kinetic/); this file is not a summary of intent, it is the rulebook the adapters must implement.

One formula per category, per-protocol data sources. Every factor's formula, scale, and thresholds are fixed here and identical across every protocol in a category — and across markets: the three Blend pools Stenion scores run one adapter and one rulebook, differing only in the pool each reads. Each category owns a section below, holding its own factor list, weight table, worked example and version changelog. There are two — lending, which every scored market runs under, and dex, published but not yet scoring anything. What legitimately differs per adapter is only where the raw inputs are read on-chain — e.g. Blend reads a per-reserve max_util cap, while Kinetic (K2), being Aave-V3-style, has no such cap and instead anchors the same utilization formula to its own OPTIMAL_UTILIZATION_RATE (see §5). The anchoring pattern ("grade against the protocol's own on-chain parameter") is the invariant; the specific parameter that pattern resolves to is a documented per-protocol fact, not a new threshold.

If the code and this document ever disagree, that is a bug — open an issue (see Disputing or changing a threshold).


Current version

Versions are per category, on independent counters that each start at 1, so a version number alone does not identify a rulebook — the category and the number together do. One row per category, and the changelog behind each lives in that category's own section:

CategoryCurrent versionScored todayRulebook and changelog
lending1yesLending
dex1yesDex

The two 1s are not the same 1. Counters are independent and each starts at 1, so dex v1 and lending v1 are two different rulebooks rather than two editions of one, and neither is older than the other. The category and the number together identify a rulebook; the number alone does not.

Dex, methodology v1 — complete, and in use since 2026-08-29. The two-factor AMM rulebook in Dex — adminKeySafety weighted 0.55 and assetControlSafety weighted 0.45, each with its formula, its thresholds and its worked example. One market is scored under it — aquarius-xlm-usdc, read by adapters/aquarius/ — and every stored dex row carries version 1. The factor set was admitted first and the weight table reviewed separately, because a weight can only ever be an unvalidated judgment call and arguing one alongside the factor set would have buried it; both are version 1, since no score was ever published under the factor set alone. A third factor, depthSafety, was proposed and deferred: Aquarius publishes no unit of value to denominate a trade size in, and every way of inventing one is either a fabricated anchor or a permanent choice made on one protocol's data. No size floor exists and none is pending — neither surviving factor is size-sensitive. It was published ahead of the adapter, which is the order ../TAXONOMY.md requires: a category's rulebook must be reviewable before an adapter is written against it.

Lending, methodology v1 — the rulebook described in the Lending section, in full, including oracleSafety scoring both price freshness and manipulation resistance (§2), the minimum-size filter §4 and §5 select reserves through (The minimum-size filter), and the market-size floor that decides whether a market is scorable at all. The floor is a precondition rather than a formula — it moves no number and did not bump this version.

The rest of this section — what bumps a version, and what a boundary means for a stored score — is policy that applies to every category, not just lending.

Versioning begins here. methodology_version = 1 is the only version lending's rulebook defines, and the only one any stored row will carry. A version 2 was briefly live in the code — stamped onto runs between 2026-08-14 11:25 and 2026-08-18 11:30 UTC, before the rulebook was flattened back to v1 — and that history is discarded rather than migrated, for the same reason the development-era history was: it was computed under a rulebook that no longer exists, and nobody was downstream of it. After that discard there is no v1-versus-v2 boundary in risk_scores, and none to look for.

Earlier development history was discarded, not migrated

Before this point Stenion accumulated a few weeks of scored runs during development, under earlier iterations of these rules. That history was deleted rather than carried forward, and this is a deliberate, recorded choice rather than a silent one. Three reasons, stated plainly:

  • It contained scores computed under two known bugs, since fixed. Those numbers were wrong under their own rulebook, not merely scored under a different one.
  • It predates the oracle robustness work (§2), so its oracleSafety values measured price age alone — a signal we now consider misleading rather than merely incomplete.
  • Nobody was downstream of it. Every row came from our own cron during development; no external consumer had been built against the API, and the only reader of the history was our own score chart. Marking a discontinuity in a dataset nobody had read would have been bookkeeping, not disclosure.

A clean history starting from a rulebook we actually stand behind is more honest than a marked-up one carrying forward numbers we know were wrong. This is the last time that reasoning applies. From here on, history is never deleted and never backfilled — the version stamp exists so a change is labeled instead.

What bumps the version, going forward

Bump when a change alters what a number means — a factor starting or stopping measuring something, a threshold's anchor changing, a re-weighting, or any formula change that moves scores for unchanged on-chain state. Don't bump for a fix that makes the implementation match the rule already documented here (the stored scores were wrong, not scored under a different rulebook — say so in the changelog instead), for adding a protocol or an adapter, or for wording, disclosure, and presentation changes. The test is simple: if comparing an old score to a new one would mislead, bump; if the old score was just incorrect under this same rulebook, don't.

Scores across a boundary are not comparable

No boundary survives in the stored data — the v2 rows above were discarded — but the machinery that marks one is live and tested, because the first real bump must be legible on the day it happens rather than built in a hurry then:

  • The indexer stamps risk_scores.methodology_version from this category's entry in METHODOLOGY_VERSIONS (core/src/category.ts) at write time, resolved from the target's own category. An adapter has no say in the version; it only declares which category it belongs to. risk_scores.category is stamped beside it, because every category's counter starts at 1 and the pair — not the integer — identifies a rulebook.
  • The score-history chart on each protocol page breaks the line at a version change rather than drawing through it, and the run list labels the break. Both paths are covered by fixture tests, since live data cannot exercise them.
  • The version is returned on every history point and on the protocol detail from GET /api/v1/protocol/:id. To check which rulebook produced a stored score, read that column — don't infer it from the date.

History is not backfilled across a bump, and cannot be — risk_scores stores only outputs (the score and the factor map), never the raw on-chain inputs a run was computed from, so no one, including us, can recompute an old row under new rules.


Ground rules (non-negotiable)

  1. The same formula applies to every protocol in a category, with no exceptions. A factor's formula, its weight, and its thresholds are fixed in that category's section below, in one shared place. They do not vary per protocol, per adapter, or per anything else within the category. A category boundary is the one and only place a rule may differ — and it differs because the factor sets differ, not because a protocol asked: utilizationSafety weighted 0.20 says nothing about an AMM with no borrow cap. That is a different rulebook, published in full under its own heading and versioned on its own counter, never lending's rules bent to fit. Two protocols in the same category are always graded by the same rules.
  2. Payment never changes a threshold or a formula. Protocols can pay for visibility, speed, or private tooling — never for a better number. A paid tier cannot move a threshold, reweight a factor, or alter a curve. The only thing that changes a protocol's output is its own real, on-chain data.
  3. Different protocols can and should score differently. That is the point. What must never differ is the rule being applied. Blend scoring 54 and a hypothetical protocol scoring 80 is a result of their data, not of two different rulebooks.
  4. No fabricated numbers. Where real data genuinely isn't available for a factor, the score uses a clearly-flagged neutral baseline (called out explicitly below) — never an invented, plausible-looking value.
  5. AI never sets a score. Any AI feature only explains or summarizes the numbers these formulas produce. It never generates an independent risk assessment.

Score model

  • Overall score: 0–100, higher = safer. API/field name safetyScore.

  • Every factor is on the same scale: 0–100, higher = safer. Factor names end in *Safety so a name never disagrees with its number — a collateralSafety of 70 means well-diversified (safe), not "70% concentrated."

  • The overall score is a weighted mean of the category's factors, renormalized over whichever factors are non-null (so a genuinely inapplicable factor doesn't drag the score toward zero rather than being excluded):

    safetyScore = round( Σ(factor.value × factor.weight) / Σ(factor.weight) )
    

This arithmetic is shared; the factors it averages are not. The formula above is category-agnostic — it reads a value and a weight and nothing else — and it is implemented once, in scoreFactors (core/src/scoring.ts), which every adapter of every category calls. Which factors exist and what each is weighted is per category: declared once in CATEGORY_FACTORS (core/src/weights.ts) and published in that category's own section below. So the weight table and the factor list live under Lending, not here — a second category would bring its own, not edit lending's.

A factor may publish a components breakdown — the sub-signals behind its value. Components with a numeric value are what the factor was computed from; components with a null value are disclosures: real, readable on-chain quantities we publish but deliberately do not grade, because scoring them would invent comparability the data does not support (see §2c and §2d). A null component is never missing data.

One published field sits outside this formula entirely: operationalState, which reports which user operations a market's own contracts are currently refusing. It is not a factor, not a multiplier, and not an input to safetyScore — see Operational state is published, never scored for why, and for why that is a decision rather than an omission.


Lending

Everything from here to the Dex heading is lending's rulebook and lending's alone — its version changelog, its factor weights, its worked example, and the five factors themselves. It is one of two categories Stenion publishes a rulebook for (PROTOCOL_CATEGORIES in core/src/category.ts), and it is the only one anything is scored under today. This section was written as the shape a second one would take — its own heading, its own changelog, its own weight table, its own factor list — and Dex is that second one, taking exactly that shape. Nothing above this line is lending-specific; nothing below it may be assumed to hold for a category that isn't lending, and in particular a dex score is not comparable with a lending one.

Which protocols are scored under it: every market on the registry today — the three Blend pools (adapters/blend/) and Kinetic/K2 (adapters/kinetic/). Ground rule 1 binds all of them to what follows.

Version changelog

Every scored run is stamped with the rulebook version that produced it (risk_scores.methodology_version, from the lending entry in METHODOLOGY_VERSIONS in core/src/category.ts), and it is surfaced on the API's protocol detail and on each history point. Versions are per category and counters are independent, so this changelog is lending's; another category's v1 is a different rulebook, not an earlier one. Each category gets exactly one such table, in its own section. The changelog:

VersionEffectiveChange
1initialThe five-factor model as documented here. oracleSafety scores price freshness and manipulation resistance (§2); liquiditySafety/utilizationSafety score only reserves clearing the minimum-size filter (§4).

One row, and that is the point: v1 is where versioning starts, not where it started counting again. Development-era history under earlier iterations of these rules was discarded rather than migrated — what that was and why it was deleted is at the top of this document, along with what does and doesn't warrant a bump. See Current version.

Corrections that did not bump the version

Fixes where the implementation disagreed with this document and the document was right. The rulebook did not change, so these are not version boundaries and stored scores remain comparable across them — but they are recorded here rather than left silent, because a score did change shape even if no published number moved.

The entry below was verified against the development-era history that has since been discarded (see Current version), so its row counts are no longer re-checkable. It is kept as the record of a correction, not as a live claim about stored data. The fix itself is in the shipped v1 rulebook.

Date (UTC)Correction
2026-08-16liquiditySafety (§4) and utilizationSafety (§5) returned 100 when no reserve qualified for their minimum, in both adapters. Both are a minimum over a filtered set of reserves; over an empty set that is undefined, not the top of the scale — so an unassessable pool published "maximally safe" from no data, contrary to ground rule 4. Both now return 0, matching collateralSafety's existing treatment of the same case. No published score was affected: the path had never executed — verified by scanning the entire stored history of both protocols for the signature the defect leaves in a factor's detail (a worst reserve (…) naming no asset, since worstAsset stayed empty when nothing was measured), with zero matches. Re-checked on 2026-08-16 against 1,923 rows; liquiditySafety has ranged 20–34 and utilizationSafety 10–18 across that history, never approaching the 100 the empty path would have published.

Amendments folded into v1

Changes to the rulebook made while v1 was still being finalized as the comparability baseline — before any surviving stored history existed. These are not version boundaries: v1 is defined as the rulebook this document describes, and these are part of that definition rather than a departure from it. None of them left a step in a published history, because the history they predate was discarded rather than carried forward. Each gets a row anyway, so that what v1 means is traceable rather than assumed: "no bump required" does not mean "no record required".

This section closes when the next change lands. From that point the rule in What bumps the version applies without exception, and a change that alters what a number means bumps to v2.

Date (UTC)Amendment
2026-08-18§4 and §5 gained the minimum-size filter: both now select the worst reserve only among reserves clearing the protocol's own declared minimum exposure or 0.5% of the pool's supplied USD. This changes what the two factors measure, which is why it is recorded rather than treated as a correction. No live score moved when it landed — verified against both protocols on the day: Blend excludes nothing (its smallest reserve is ~$3.4M against a $5.00 min_collateral), and K2's dust reserve had already stopped being its worst reserve. What it changes, measured on the frozen 2026-08-16 snapshot: liquiditySafety 34 → 44 and utilizationSafety 18 → 30 (score 24 → 28), by excluding a $3.00 reserve holding 0.19% of a $1,571 pool.

History is not backfilled across a version bump, and cannot be. risk_scores stores only outputs — the score and the factor map — never the raw on-chain inputs a run was computed from, so an old row cannot be recomputed under new rules by us or by anyone. The discontinuity is real and permanent; the version stamp exists so it is legible rather than appearing as an unexplained step in a chart.

Factor weights

Lending's weights, and lending's only. This table is the published face of CATEGORY_FACTORS.lending in core/src/weights.ts, which is where the adapters read them from — neither adapter contains a weight of its own, and core/src/scoring.test.ts parses this table and fails if the two disagree in either direction.

FactorWeight
oracleSafety0.25
collateralSafety0.20
adminKeySafety0.20
utilizationSafety0.20
liquiditySafety0.15
Total1.00

Worked example (live Blend Fixed V2 pool, 2026-08-14, lending methodology v1): 70×0.20 + 100×0.25 + 40×0.20 + 22×0.15 + 14×0.20 = 53.1 → 53.

Weights are an unvalidated judgment call, not an external fact. oracleSafety carries the most weight because an untrustworthy price silently poisons every other measurement — collateral value, utilization and liquidity are all priced off it. Liquidity carries the least because it partly overlaps utilization. There is no external framework these exact weights are anchored to yet — they are open to challenge like any threshold below.

Oracle robustness was folded into oracleSafety rather than given its own factor, partly for this reason: a sixth member would have forced a redistribution across all five, layering a second unanchored judgment call on top of one already flagged as unanchored. The taxonomy in core/src/types.ts stays at five factors.


The five lending factors

For each factor: the exact raw on-chain data that feeds it, the exact formula, and why the thresholds are what they are (anchored to an external/on-chain value where one exists, labeled an unvalidated judgment call where none does).

Two fixed-point scalars appear throughout, taken from blend-contracts-v2/pool/src/constants.rs:

  • SCALAR_7 = 10^7 — decimals for c_factor, l_factor, util, max_util.
  • SCALAR_12 = 10^12 — decimals for d_rate, b_rate.

A reserve's human-unit totals (used by several factors) are:

supplied = b_supply × b_rate / (SCALAR_12 × 10^assetDecimals)
borrowed = d_supply × d_rate / (SCALAR_12 × 10^assetDecimals)

1. collateralSafety — collateral concentration (weight 0.20)

What it measures: how spread out the pool's supplied value is across its reserves. A pool whose value sits in one asset is far more exposed to a single de-peg or liquidation cascade than a balanced one.

Raw on-chain data (Soroban RPC, no third party):

  • Per reserve, from the pool contract's persistent storage:
    • ResData entry → b_supply, b_rate
    • ResConfig entry → decimals
  • Oracle price per asset: lastprice(Asset::Stellar(address)) on the pool's configured oracle contract → price, and the oracle's decimals().
  • USD value per reserve: suppliedUsd = supplied × (price / 10^oracleDecimals).

Formula — a normalized Herfindahl–Hirschman Index (HHI) over each reserve's share of total supplied USD:

Let vᵢ = supplied USD of reserve i (only priced reserves with vᵢ > 0)
    n  = number of such reserves
    sᵢ = vᵢ / Σv                 (each reserve's share)
    HHI = Σ sᵢ²                   (ranges from 1/n for a perfectly even split, to 1)

collateralSafety = clamp( (1 − HHI) / (1 − 1/n) × 100 , 0, 100 )

Edge cases: 0 priced reserves → 0 (can't assess, treated as unsafe rather than guessed); exactly 1 priced reserve → 0 (fully concentrated by definition).

Why HHI / why these anchors: HHI is the standard, widely-published concentration measure (used by competition regulators and in portfolio analysis) — an external framework rather than a Stenion invention. The anchoring points are not arbitrary either: 1/n (a perfectly even split) is the mathematically safest achievable state for n reserves and maps to 100; 1 (everything in one asset) is the worst and maps to 0. Normalizing by 1/n means the score grades a pool against the best it could do given how many reserves it has, not against an arbitrary constant.


2. oracleSafety — price trustworthiness: freshness and manipulation resistance (weight 0.25)

Why this factor is not just price age. An age-only oracle factor scores a fresh but manipulated price 100 — which is precisely the configuration behind the February 2026 YieldBlox/Blend incident. Freshness alone is not a weak signal on that axis, it is a misleading one, so this factor takes the binding constraint of freshness and manipulation resistance. Earlier development-era scores did measure age alone; that history was discarded rather than carried forward, and no stored row was computed that way — see Current version.

What it measures: whether the prices this pool actually runs on can be trusted. Two things must both hold, and the factor takes the binding constraint of the two — a bounded stale price and a fresh unbounded price are both untrustworthy, for different reasons:

oracleSafety = min( priceFreshness , deviationBound )

Both sub-signals take the worst reserve, the same convention as every other factor: the binding constraint is the single weakest reserve, and averaging would hide it. Both are published in the factor's components array so the composite is never an opaque number.

Both are also anchored to parameters the pool's own price path has to publish, which makes this the one factor with a precondition attached: a market whose oracle publishes neither a staleness tolerance nor a deviation bound is not scored at all, rather than scored with a guessed anchor or with this factor dropped. See 2e, the oracle-legibility precondition.

Every reserve at the binding value is named, not one of them. When several reserves tie on a sub-signal the detail lists all of them; when all of them tie it says so rather than singling one out. This is reporting only — the published value is the same minimum either way — but it is load-bearing for reading a score honestly. Blend prices its whole pool from one aggregator publish round, so its reserves carry identical ages and always tie on freshness. Naming one of them would make an iteration-order artifact read as a diagnosis, and a reserve name that is really a tie-break is worse than no name at all.

2a. priceFreshness — how stale the worst price is

Raw on-chain data (Soroban RPC): per reserve, the price's publish timestamp from the method the protocol's own pool calls; fetchedAt is the adapter's read time; age = fetchedAt − timestamp.

fresh = the protocol's own publish/refresh interval        → 100
dead  = min( protocol's own max acceptable price age, 3600s ) → 0

priceFreshness = clamp( (age − dead) / (fresh − dead) × 100 , 0, 100 )

No usable price for a reserve → 0 (a missing feed is maximally unsafe, not skipped).

Both anchors are the protocol's own on-chain parameters, the same anchoring pattern utilizationSafety uses. Which parameter each resolves to is a documented per-protocol fact, not a per-protocol rule:

Protocolfresh sourcedead source
Blendoracles()[i].resolution on the pool's oracle aggregator (300s)max_age() on the same aggregator (900s)
Kinetic (K2)PriceCacheTtl on the price oracle (30s) — the window inside which K2 itself treats a price as current; K2 exposes no publish-interval getterthe tighter of the per-asset max_age (43200s) and the global price_staleness_threshold (3600s)

Taking the tighter of two limits a protocol declared is not a Stenion threshold — both numbers are K2's, and the binding one is the one that governs.

⚠️ The 3600s cap on dead is the one Stenion constant left in this factor, and it is an unvalidated judgment call. Anchoring purely to a protocol's own max age would mean a protocol scores better for tolerating staler prices — K2's per-asset max_age is 12 hours, which would make a six-hour-old price score ~50. That is the wrong incentive for a platform protocols are ranked by, so the anchor is capped. There is no external framework fixing the cap at one hour; it is open to challenge like any threshold here. It lives in one place, STALE_CEILING_SECONDS in core/src/scoring.ts.

2b. deviationBound — can a single update move the price arbitrarily far?

Binary, not a curve:

deviationBound = 100  if the pool's price path bounds a single-step move, and that bound is armed
                 0    otherwise
ProtocolBounded when
Blendthe aggregator's per-asset max_dev satisfies 0 < max_dev < 100 — the contract's own condition in oracle-aggregator/src/price_data.rs
Kinetic (K2)max_price_change_bps > 0 and get_last_price(asset) returns a present, non-zero baseline

Why the extra clause for K2. The two contracts fail in opposite directions when there is no prior price to compare against. Blend's aggregator fails closed: with no older record it returns None and the reserve simply cannot be priced. K2's validate_price_change fails open: with no stored baseline it returns Ok and lets any price through, so a configured bound with no baseline is inert. Checking the baseline is what distinguishes a breaker that is configured from one that is actually armed — and the no-baseline case is exactly the newly-listed-thin-asset scenario that the YieldBlox incident ran through.

Why this is anchored, and what it isn't. The scored quantity is the presence and arming of a bound, and it is read from the protocol's own on-chain configuration — the same pattern as utilizationSafety's max_util. max_dev = 0 does not mean "a tight bound of zero"; the aggregator's own type documentation says "If this is 0, the oracle will just fetch the last price within the resolution time" — the check is skipped entirely. That is provably the condition that permits an unbounded single-step move.

Base assets are excluded, not scored 0. The Blend aggregator's lastprice short-circuits its Base and BaseAssets to exactly 1.0 at the current ledger time without consulting any upstream feed. Those reserves have no oracle-derived price to grade, so they are dropped from both sub-signals and the count of excluded assets is disclosed. (Whether such a peg holds is a real risk — but it is a collateral/peg question, not an oracle-robustness one, and inventing a number for it here would be the kind of fabrication ground rule 4 forbids.)

2c. Per-feed price ages are disclosed, never scored

priceFreshness grades the worst reserve, which is the right thing to score but hides the spread — and on real data the spread is the informative part. A factor value of 0 reads as a general condition of the oracle. "Two feeds have not updated in hours while two others update every few seconds, through one contract and one source" is a specific, checkable statement about which feeds are being maintained, and it is the one a depositor can act on.

So every reserve's price age is published as a disclosure-only component (priceAges, value: null), ordered oldest-first, alongside a count of how many exceed the protocol's own declared staleness limit — Blend's aggregator max_age, K2's price_staleness_threshold. The count is therefore a statement about the protocol's own rules, not about a Stenion line. A reserve with no usable price at all sorts as the oldest rather than the freshest.

It is not scored, because it would double-count: these are the same ages priceFreshness was computed from, republished so that the grading can be checked rather than taken on faith. It is published on healthy pools as well as unhealthy ones — a disclosure that appears only where trouble is expected gives a reader no baseline to compare against.

2d. Bound tightness is disclosed, never scored

The raw bound is published as a disclosure-only component (value: null) — visible, never graded. Grading it would invent comparability the underlying data does not support:

Blend max_devK2 max_price_change_bps
Scopeper assetglobal
Unitswhole percent (60 = 60%)basis points (2000 = 20%)
Baselinethe previous upstream record, one resolution step backget_last_price — the last price the contract served
Bounds move perpublish interval (300s)query — no fixed time spacing

Both compute |new − old| / old, and the unit difference normalizes trivially. The baseline difference does not: "20% per arbitrary interval" and "60% per five minutes" are different quantities, so the intuitive reading that K2's bound is three times tighter than Blend's is unsound. Publishing the numbers side by side without a score is the honest treatment.

2e. The oracle-legibility precondition

Both halves of this factor are anchored to parameters the pool's own price path publishes: §2a's window comes from resolution and max_age, §2b's bound from per-asset max_dev. That is the whole design — the numbers are the protocol's, not Stenion's. It has a precondition hiding inside it, which this section makes explicit:

A market is scorable only if its price path publishes the parameters §2 grades against. Where it does not, the market is not scored at all — it is not scored with a guessed anchor, and it is not scored with oracleSafety omitted.

This is the market-size floor one factor down, and it has the same shape: a precondition on what gets scored at all, rather than a rule about how to score it. Markets excluded by it are published in dashboard/app/lib/coverage.ts as oracle-not-gradable, with a per-market reason and a date — never as a protocols row, and never with a numeral.

Where the line falls today. Blend's oracle-aggregator publishes all three reads (max_age(), oracles(), asset_configs()); those are Blend's interface, not SEP-40's. Every Blend V2 pool Stenion scores sits on one. Four live pools do not, and the interfaces below were read on 2026-08-26 out of each oracle's own wasm (contractspecv0 via Soroban RPC getContractMethods) rather than probed by calling a list of guessed names — a guess-list cannot distinguish "this contract lacks the method" from "this is a different contract":

PoolOracle wasmWhat the contract actually ismax_ageoraclesasset_configs
Blend Fixed41df0489…Blend oracle-aggregator✅✅✅
YieldBlox8cf43882…oracle-aggregator, different build✅✅✅
Etherfuse65300c00…oracle-aggregator✅✅✅
Orbita71a844e…bridge oracle — ctor (admin, stellar_oracle, other_oracle)~~~
Forex1d1c90d3…proxy — CONFIG.base_oracle points one hop up at a SEP-40 feed~~~
Spectra PTs4a444181…deterministic zero-coupon-bond pricer — not a feed at all~~~
Solv5700be21…SEP-40 feed registry — the only SEP-40 implementer of the four~~~

Those four are not one shape. They are four different contracts with four different wasm hashes doing four different things, and the only thing they agree on is the column that decides this: none answers any of the three. Worth stating because the obvious fix — "handle the other oracle shape too" — is really "handle four more shapes", and each would be a separate reading of a separate contract's semantics. All four answer decimals() and lastprice(), so §1, §3, §4 and §5 compute normally for them; what is missing is only the metadata §2 grades against.

Why there is no fallback anchor: SEP-40 does not define one

Stated plainly rather than worked around. SEP-40 defines base, assets, decimals, resolution, price, prices, lastprice — no maximum acceptable price age and no deviation bound anywhere in the interface. The spec puts staleness checking on the consumer ("Always check retrieved price data for staleness by comparing the quoted timestamp with current date"), which is precisely the judgment §2 exists to make and precisely what it refuses to make from an invented number.

So the only candidate is resolution() — a publish interval, not a staleness tolerance — and using it fabricates a 100 on a demonstrably stale price:

Solv publishes resolution() = 43200 (12 hours, and mutable after deployment via its own set_resolution, which its source documents as a deliberate deviation from SEP-40). Fed to freshnessWindow with no max_age to pair it with, that yields {fresh: 43200, dead: 86400} — because STALE_CEILING_SECONDS clamps dead and not fresh, so a feed that declares a slow tick gets a slow dead line rather than a capped one. Solv's genuinely stale feeds, read at 10,285s and 21,739s old on 2026-08-26, would both publish priceFreshness 100.

And resolution() is only present on one of the four at all. Orbit and Forex expose it one hop upstream through their bridge/proxy, which would make the anchor a property of a contract the pool does not itself publish; Spectra has no upstream feed to chase.

Two of the four price off the ledger clock, so freshness is 100 by construction

The sharper problem, and the reason this is a precondition rather than a "weak signal":

  • Spectra PTs runs "Spectra Deterministic Oracle — Zero Coupon Bond Model". Its price is a function of start_t, maturity and initial_implied_apy evaluated at the current ledger time, its lastprice ignores the asset argument entirely (its own doc comment says so), and its owner can move the target with set_future_pt_value. There is no publish event, so there is no such thing as a stale price: age is always ~0 and any freshness formula returns 100 permanently, whatever happens to the asset.
  • Orbit's dominant reserve — 99.5% of that pool's $190,863 on 2026-08-26 — returns exactly 1.0 at current ledger time while touching no upstream contract. §2b already handles that case on an aggregator: base assets are excluded rather than scored, because there is no oracle-derived price to grade. Orbit's bridge publishes no base(), so there is nothing to detect them with, and the reserve would be graded as a fresh feed.

A fabricated 100 is worse than a fabricated 0, because 0 at least renders in the danger band where a reader discounts it. Ground rule 4 forbids both.

What was rejected, and why it must not be re-proposed
  • Make the three calls optional and score oracleSafety on what remains. Rejected: there is nothing to score on. Both anchors are gone, deviationBound collapses to a constant 0 for all four pools — a constant is not a measurement — and priceFreshness collapses to a constant 100 for two of them. It also reads as a diagnosis it did not make: a max_dev of 0 means a protocol disabled its bound, which is the YieldBlox finding; publishing the same 0 for a protocol whose oracle never had the mechanism asserts a choice nobody made.

  • A null oracleSafety, scored on the remaining four factors. Rejected, and this is the one that looks reasonable until it is measured. scoreFactors renormalizes over non-null weights, so dropping oracleSafety divides by 0.75 instead of 1.00. Run against live chain data on 2026-08-26, that is not neutral — it is a large upward revision, because oracleSafety is the heaviest factor (0.25) and the one a badly-configured pool scores worst on:

    PoolScore with oracleSafety: nullFor comparison, live registry
    Orbit71Blend 51
    Spectra PTs49Kinetic 27
    Solv15YieldBlox 25
    Forex11

    Orbit would publish 71 — the highest number in the registry, twenty points above Blend Fixed — while Admin-Frozen, 99.5% concentrated in a synthetic priced at a hardcoded 1.0, through an oracle publishing neither a staleness tolerance nor a deviation bound. That is an incentive inversion: YieldBlox scores oracleSafety 0 for having a deviation bound and disabling it, so not publishing the mechanism at all would pay better than publishing it switched off. Ship an opaque oracle, get a better number — which is ground rule 2's concern arriving through the back door.

    It is not the same failure as the rejected sixth factor under Factor weights: that one moved every protocol's score by changing the denominator for everyone. This moves only the affected pool's. The failure here is different and, in a ranked list, worse — two entries in the same ranked column would be graded on different rulebooks, four factors against five, which ground rule 1 does not permit. A null factor is defined as "genuinely doesn't apply to this protocol" (the dashboard renders it "Not applicable to this protocol"). A pool that runs on prices, whose price configuration we could not read, is not a pool to which price trustworthiness does not apply.

  • Reading the anchor from the upstream oracle a bridge or proxy forwards to. Rejected: it answers for a contract the pool does not publish and does not itself constrain, it exists for only two of the four, and following it is a bespoke traversal per oracle implementation — a per-market rulebook in all but name. It also would not have helped: the upstream feeds behind Orbit and Forex publish resolution() and no max_age, so the traversal lands back on the fabricated anchor above.

This does not bump lending's methodology version. It moves no published number and changes no formula. Every market Stenion scores runs on an aggregator, so oracleSafety is computed byte-identically before and after; what changed is that a precondition already implicit in §2 is now written down and enforced in code (ORACLE_GRADING_READS and oracleNotGradable), instead of surfacing as an unexplained HostError from whichever read happened to run first. Same reasoning, and the same conclusion, as the market-size floor.

⚠️ What this costs, stated rather than glossed. Four live markets stay unranked, one of them holding real money — Orbit's $190,863 on 2026-08-26. The market-size floor's warning applies here in full: an excluded market has no entry at all, which is a stronger action than excluding a reserve. coverage.ts is the answer to that and not a formality — each of the four gets a page, a per-market reason, the contract addresses, and a verify sentence, so a reader can disagree with the decision from the same data it was made on. What they do not get is a number, because there is no number to give them.

The precondition is a property of the oracle, not a verdict on the protocol. Nothing here says these markets are unsafe, or that their oracles are bad ones. A deterministic bond pricer is a perfectly coherent way to price a principal token; it is simply not a thing this factor knows how to grade. If such an oracle later publishes a staleness tolerance and a deviation bound, the pool becomes scorable with no rule change — it is a BLEND_POOLS entry and a deleted coverage entry, in one PR.

What was considered and deliberately rejected

Recorded so these are not re-proposed as improvements later. Each was investigated against the February 2026 YieldBlox incident — the test being whether it would have distinguished the manipulated price from a legitimate one at the time, since a signal that looks sophisticated but would not have caught the actual attack is worse than none: it manufactures confidence.

  • Filtering oracleSafety by reserve size, the way §4/§5 are filtered. Rejected on principle, not on impact — and the distinction it turns on is the reason §4/§5 may be size-filtered while this factor may not:

    §4 and §5 measure current state. §2 measures a vulnerability. How drained a reserve is right now means little when the reserve holds $4, because the exposure is capped by what is actually in there. Whether a price can be trusted is not capped that way, because the attacker's move is to grow a position against the mispriced asset. A dust reserve with a stale price is an open door, not a small room. Its balance today says nothing about what can be borrowed against it tomorrow.

    And it would blind the factor to the exact scenario it exists for: a newly-listed thin asset with a bad price is the shape the February 2026 YieldBlox incident ran through, which §2b already names. A filter that removes thin assets from an oracle-trust factor removes the attack it was built to catch.

  • A Stenion-computed deviation from the oracle's price history (calling Reflector's prices(asset, N) ourselves and comparing the latest price to a trailing mean). Rejected — this would have made the platform actively worse. It is a coincident indicator, not a leading one: it can only fire while an attack is in progress, and only if the indexer happens to sample inside the manipulation window. The indexer runs every five minutes, so the overwhelmingly likely outcome is that it reads clean and Stenion publishes a confident oracleSafety of 100 during an active exploit. It also measures a code path the pools never consult: neither Blend's aggregator nor K2's oracle exposes price history to the pool at all. A signal that is usually silent during the event it claims to detect, computed over data the protocol does not use, is not a weak signal — it is a misleading one.

  • TWAP. Not available: the deployed Reflector contracts (version() == 6) expose no twap method — the exported interface is base, assets, decimals, resolution, price, prices, lastprice, last_timestamp, history_retention_period, …. Earlier Reflector versions had one; the live contracts do not. Neither protocol's oracle passes history through either. And on the merits it would not have helped: the attacker held the only trades in the window, so a short TWAP over a dead order book is the manipulated price.

  • Oracle type / provider identity ("is it Reflector?"). Zero discriminating power: the exploited pool and the healthy Blend pool both price through Reflector-family feeds via the same oracle-aggregator contract family. What differed was configuration, not provider.

  • Number of upstream sources. Both pools had exactly one upstream oracle, so it would not have separated them. It is also not comparable across protocols: a count is only readable where a contract happens to publish one (K2's upstream RedStone adapter exposes unique_signer_threshold() == 3; Reflector's node consensus is not exposed on-chain at all), so counting would systematically understate feeds that keep their aggregation internal.

  • SDEX order-book depth via Horizon. Conceptually the right quantity — thin market depth is what made the manipulation cheap — and mechanically readable. Rejected for now on three grounds: order books are trivially spoofable with walls that are never hit; it only applies to assets priced off the Stellar DEX; and it cannot be validated retroactively, because the exploited market has since been rebuilt. Tracked as a candidate in ROADMAP.md rather than shipped on intuition.

What this factor would have said on 2026-02-22

Running the shipped adapter against the exploited pool and the healthy one today, same rulebook, no special-casing:

PoolpriceFreshnessdeviationBoundoracleSafety
Blend Fixed V2 (CAJJZSGM…)100100 — all reserves bounded (max_dev 60/20/20)100
YieldBlox (CCCCIQSD…, the exploited pool)840 — XLM and AQUA carry max_dev: 0, check disabled0

Both pools' prices are fresh, so an age-only factor scores both high — 100 and 84, the latter being an ordinary mid-window price age, not a warning. This factor separates them anyway, and on the axis that actually failed.

This is no longer a demonstration run. As of the multi-pool change, the YieldBlox pool is a registered, continuously scored entry in the public registry, and the row above is its live oracleSafety, published every five minutes like any other. Two consequences worth stating: the claim in this section is now checkable by anyone against GET /api/v1/protocol/yieldblox rather than reproducible only by running the adapter by hand; and the number will move, because it is live. The pairing that matters — a fresh price and a disabled bound — is a property of the pool's configuration, not of the moment it was sampled.

The entry is labelled a Blend V2 pool wherever it appears (deployedOn on both API responses). It is not a third protocol, and the registry must not be read as saying so.

⚠️ Two honest limits on that claim, stated rather than glossed:

  1. The historical max_dev is a deduction, not a reading. Soroban RPC serves no historical contract state, so the exact value USTRY carried on 2026-02-22 cannot be read back. What is verifiable: the deployed aggregator skips the deviation check entirely when max_dev is 0 or ≥ 100, and rejects the price outright otherwise — so a ~100× single-step move is arithmetically incapable of passing any bound between 1 and 99. USTRY's bound must therefore have been disabled. USTRY today carries max_dev: 10; XLM and AQUA in that same live contract still carry 0.
  2. Semantics were verified against the public repo, not that binary. The exploited pool's aggregator (wasm 8cf43882…) and Blend Fixed V2's (41df0489…) are different builds. Both export the same eleven functions, and the max_dev logic above is read from blend-capital/oracle-aggregator; it has not been decompiled from the exploited pool's specific binary.

⚠️ K2's enforcement is an inference, held to the same standard. max_price_change_bps is enforced on every return path of get_asset_price_data in K2's audited source (code-423n4/2026-04-k2), and the deployed wasm contains both the max_price_change_bps and PriceChangeTooLarge symbols. But the live kinetic_router does not call that method — it calls get_asset_prices_vec_fresh, one of nine functions present in the deployed oracle and absent from the audited source, whose source is not public. The audited sibling get_asset_prices_vec does enforce the breaker, and all three methods return identical data today. We score it as enforced on that basis. That is an inference, not a verification, and it is written up as a finding in its own right — see the Kinetic entry in the registry.


3. adminKeySafety — admin signer structure + activity (weight 0.20)

What it measures: how much unilateral, live control a single party has over the pool. A lone hot key that can reconfigure the pool is the sharpest centralization risk; multisig and inactivity are safer.

Raw on-chain data:

  • The admin address comes from the pool contract's instance storage (Admin, or Config.admin) via Soroban RPC.
  • If the admin is a keypair account (G…), signer structure and activity come from Horizon (official Stellar infra, not a third party):
    • GET /accounts/{address} → thresholds.high_threshold, signers[] (→ signerCount)
    • GET /accounts/{address}/operations?order=desc&limit=200 → created_at of each op; recentOps = count within the last 30 days.
  • If the admin is a contract (C…), Horizon has no account entry to introspect — there is genuinely nothing to measure.

Formula — a tiered base (categorical, NOT a curve) minus a continuous activity penalty:

This factor is deliberately tiered, not a continuous function, because signer structure is categorical. The base value is chosen by tier:

TierBaseDetected by
Contract-governed admin60admin address starts with C… (flagged neutral baseline — see below)
Single master key40keypair account, not multisig
N-of-M multisig (N ≥ 2)90signerCount > 1 AND high_threshold > 1
Multisig + timelock100RESERVED — see note

Then a continuous activity penalty is subtracted:

activityPenalty = min(30, recentOps × 3)          # capped so structure still dominates
adminKeySafety  = clamp( base − activityPenalty , 0, 100 )

⚠️ The "Multisig + timelock" (100) tier is reserved and not yet reachable. No on-chain timelock signal is exposed to the adapter through Horizon today, so nothing is ever scored 100 by this factor at present. It is documented as the intended top tier so that when a timelock signal becomes detectable, the tier already exists rather than being invented ad hoc. This is an aspirational placeholder, explicitly flagged, not a live rule.

⚠️ The contract-governed baseline (60) is a flagged neutral value, not a measurement. When the admin is a contract, we cannot introspect its governance via Horizon. Rather than fabricate a plausible signer/activity number, we assign a fixed, clearly-labeled neutral baseline of 60 and say so in the factor's detail string. This is honest ignorance, not a score.

Why these numbers (unvalidated judgment calls, partially anchored):

  • The single-key (40) vs multisig (90) split is anchored to a real, hard security fact: a 1-of-1 key is a single point of unilateral compromise; an N-of-M multisig with high_threshold > 1 provably requires more than one party to reconfigure the pool. The detection condition (signerCount > 1 AND high_threshold > 1) reads Stellar's actual account threshold model, not a proxy.
  • The exact base values (40, 90, 60) and the activity penalty shape (−3 per op, capped at −30) are unvalidated judgment calls. The cap deliberately keeps structure dominant over activity (a busy multisig should still beat an idle single key). There is no external framework these specific integers are anchored to — they are open to challenge.

4. liquiditySafety — free-liquidity depth (weight 0.15)

What it measures: the absolute withdrawal/liquidation cushion — how much value could leave before the pool is drained. Distinct from utilizationSafety, which measures proximity to the configured cap rather than absolute headroom.

Raw on-chain data (Soroban RPC): per reserve, ResData (b_supply, b_rate, d_supply, d_rate) and ResConfig (decimals), used to compute supplied and borrowed per the totals formula above.

Formula — free-liquidity share of the worst reserve, over reserves large enough to be scored (see The minimum-size filter below):

For each reserve with supplied > 0 that passes the minimum-size filter:
    free = clamp( (supplied − borrowed) / supplied × 100 , 0, 100 )

liquiditySafety = min(free) across all such reserves     # worst reserve wins

Edge cases, both → 0: no reserve with supplied > 0, and every reserve excluded by the minimum-size filter. A minimum over an empty set is undefined, not the top of the scale — an unassessable pool is reported as unassessable, the same way §1 treats having nothing to price. Returning 100 here would publish "maximally safe" derived from no data, which ground rule 4 forbids. The two are reported with different detail strings: "the pool is empty" and "everything in it is too small to grade" are different findings.


The minimum-size filter

Applies to liquiditySafety (§4) and utilizationSafety (§5) only, identically for every protocol.

The problem it solves. Both factors select the worst reserve, so a reserve holding effectively nothing can set a protocol's published number. On the 2026-08-16 Kinetic snapshot a $3.00 PYUSD reserve — 0.19% of a $1,571 pool — was the worst reserve on both factors and set liquiditySafety to 34 and utilizationSafety to 18. Nobody's capital was meaningfully exposed to it. That is a misleading number, not a conservative one.

The rule. A reserve is scored if either test passes, and excluded only when both fail:

LegTestAnchor
AsuppliedUsd ≥ the protocol's own declared minimum viable exposurethe protocol's own on-chain parameter, where it declares one
BsuppliedUsd ≥ 0.5% of the pool's own total supplied USDnone — an unvalidated judgment call (see below)

Leg A is per-protocol in exactly the sense §5's cap is: the pattern ("grade against a parameter the protocol set itself") is the invariant, and which parameter it resolves to is a documented per-protocol fact.

Protocol / marketLeg A sourceValue
Blend — Fixed V2PoolConfig.min_collateral, read live from pool instance storage, denominated in the oracle's base asset (Other:USD, 7 decimals)50000000 = $5.00
Blend — YieldBlox V2the same field, read live from this pool's own instance storage — read per pool, never inherited from the flagship50000000 = $5.00
Kinetic (K2)none — K2 declares no minimum-exposure parameter on chain. Leg B alone applies.n/a

Both live Blend pools happen to declare the same floor. That is a coincidence of their configuration, not a property of the adapter: leg A is resolved from whichever pool an adapter instance was pointed at, and a Blend pool declaring a different min_collateral would be graded against its own.

min_collateral is Blend's own dust guard: the smallest collateral a position may hold and still borrow, set where liquidating a position stops being economically worthwhile. A reserve whose entire supplied value sits below it cannot host even one position the protocol itself considers viable. That is the same question this filter asks, which is why it is borrowed rather than invented.

K2's absence is verified, not assumed. The router's instance storage and every reserve's ReserveConfiguration bitmap were read looking for an equivalent. What K2 exposes is MINSWAP (a slippage bound), FLPREMMAX, HFLIQTH/PLIQHF (health-factor lines) and a supply/borrow cap pair in data_high — all maxima or unrelated. If K2 ever ships a minimum, leg A turns on for it with no rule change.

Why both legs, and not one. Each covers a failure the other has, both demonstrated on live data:

  • Absolute-only breaks a small pool. Any floor sized for a real market ($1k, $10k) excludes all four of K2's reserves — its entire pool is ~$1,500. Both factors would go to cannot-assess and K2's score would drop, from 28 to 15. Worse than the problem.
  • Relative-only breaks a large pool. 0.5% of Blend's $186M is ~$928,000, so a reserve holding half a million dollars of real capital would be silently dropped. Leg A keeps it at $5.

The 0.5% in leg B is an unvalidated judgment call. There is no external or on-chain framework fixing it; with STALE_CEILING_SECONDS (§2) it is one of only two Stenion-chosen constants left in the continuous factors, and it is open to challenge like any threshold here.

It is deliberately set at the low end of the band that works, because the two directions of error are not symmetric. Too low leaves a dust reserve in, which reports a misleading number. Too high excludes a small but genuinely-used reserve, which hides real risk — strictly worse. 0.25% would have flipped on the live K2 reserve between two consecutive days ($3.00, then $4.00, against a $3.85 line); 0.5% clears it both times with margin.

Excluded reserves are disclosed, never silently dropped. Each affected factor publishes an excludedReserves component with a null value — the same "measured, shown, deliberately not graded" form as §2c/§2d — naming each excluded reserve, its supplied USD, its share of the pool, and the score it would have contributed. A reader can therefore see the number the filter suppressed and disagree with the exclusion, instead of never learning of it.

⚠️ This filter gives §4 and §5 an oracle dependency they did not previously have. Both are otherwise pure balance ratios that need no price at all; the filter is USD-denominated. When no reserve can be priced, the filter does not run and every reserve is scored — the two factors degrade to exactly their pre-filter behaviour rather than refusing to score. That is the right fallback, but it means a pool's liquidity and utilization numbers mean something slightly different during an oracle outage: they are unfiltered, and a dust reserve can bind them again. An individual unpriced reserve is likewise kept, never read as worthless — "could not measure" is not "empty".

The filter cannot empty the scored set on a real pool. Shares sum to 1, so the largest reserve always holds at least 1/n, which clears 0.5% for any n ≤ 200. The all-excluded branch above is therefore unreachable in practice — it is implemented and tested synthetically anyway, because that is precisely where a "cannot assess" could quietly become a 100 again.

⚠️ OPEN QUESTION, raised by the YieldBlox pool and deliberately not resolved here. The two legs are OR'd, so leg A can override leg B — and on a small Blend pool it overrides it almost entirely. YieldBlox holds ~$1.28M, putting leg B's 0.5% line at ~$6,396; six of its eight reserves fall below that line ($39.47 to $4,243.73) and every one is scored anyway, because Blend's $5 min_collateral passes for all of them. The result is that liquiditySafety (10) and utilizationSafety (0) are both set by a reserve holding $1,096.85 — 0.086% of the pool.

That is the shape of the problem this filter was added for. On Blend's Fixed pool it is invisible: leg A is a documented no-op there, because the smallest reserve holds $3.4M. On a pool three orders of magnitude smaller, the same $5 floor is doing all the work and leg B's guard never engages.

It is recorded, not fixed. Changing it — sizing leg A relative to the pool, capping it, or making the legs AND rather than OR below some pool size — moves published numbers on a live entry, and is a threshold change under the same review bar as any other (see Disputing or changing a threshold). It is equally arguable that the current behaviour is correct: min_collateral is the pool's own statement of the smallest position worth liquidating, and a $1,097 reserve at 90% utilization is a real reserve with real depositors, not the $3.00 dust the filter was built to exclude. What is not defensible is leaving the tension undocumented, which is why it is written down here.

Why this shape / this anchor: (supplied − borrowed) / supplied is 1 − utilization, i.e. the fraction of supplied value that is actually withdrawable right now. That is a direct on-chain quantity, not a modeled one — the anchor is the pool's own balances. Taking the worst reserve rather than a pool-wide average is deliberate: liquidity crises happen in the single most-drained reserve, and averaging would hide it. The mapping (free % → score %) is 1:1 and intentionally has no free parameters to tune, so there is nothing arbitrary to anchor.


The market-size floor

The minimum-size filter one level up. That filter asks whether a reserve is big enough for its number to mean anything; this asks the same of a whole market. They are two halves of one idea — a size below which a published number stops carrying information — and they are written together so neither looks like an afterthought.

The problem it solves. K2 deploys its markets as separate router contracts running identical code, the same way Blend's factory deploys pools. Three are live on mainnet as of 2026-08-20, and two of them are empty:

MarketReservesTotal priced supplied value
K2 primary (CCTUJZLY…)USDC, XLM, PYUSD, SolvBTC$1,781
K2 SolvBTC/xSolvBTC iso (CCGXGXIL…)SolvBTC, xSolvBTC$3.62
K2 Earn / earnUSDC (CDWPVHKB…)USDC, earnUSDC$0.00

Point the shipped rulebook at either of the bottom two and it does not fail — it returns a score. Every factor falls to its can't-assess branch, and every one of those branches is 0. So a market holding nothing publishes 0, in the danger band, reading "this is dangerous" to anyone scanning the registry when what is true is "there is nothing in here." That is the same misleading-number failure §4 and §5 were filtered for, one level up, and it is worse at this level: a filtered reserve still leaves a scored market with a disclosure beside it, whereas this is the market's entire published number.

The rule. A market is scorable only if it can hold at least one position the protocol itself considers viable:

LegTestAnchor
Atotal priced supplied USD ≥ the protocol's own declared minimum viable positionthe protocol's own on-chain parameter, where it declares one
Bnone — no relative leg exists at this scale. See below; this is a real gap, not an omission—
Protocol / marketLeg A sourceValue
Blend (both pools)PoolConfig.min_collateral, read per pool50000000 = $5.00
Kinetic (K2)none declared on chain — the $5.00 above is borrowed as an analogue, and is a flagged judgment call for K2, not an anchor$5.00

This is the same parameter, and the same reasoning, that §4/§5's leg A already uses: min_collateral is the protocol's own statement of the smallest collateral a position may hold and still borrow. A market whose entire supplied value sits below it cannot host even one position the protocol itself would let borrow. There is nothing there to assess, and the number is not a measurement of risk — it is a measurement of absence.

Against the table above: K2 Earn fails on total supplied value of exactly zero. The SolvBTC/xSolvBTC market fails at $3.62. K2's primary market clears by three orders of magnitude, as do both Blend pools.

Why there is no relative leg, unlike §4/§5. The reserve filter has two legs because each covers a failure the other has. No such second leg exists here. Relative to the market's own reserves is what §4 and §5 already do. Relative to the other markets in the registry would make one market's listing depend on another market's size — a market could become unlistable because a different one grew, while nothing about its own on-chain state changed. A rule about a market's own data must not have that property. So this floor is absolute-only, which is precisely the shape §4/§5 rejected as insufficient on its own, and that limitation is the reason it is set low rather than at a number that sounds meaningful.

The direction of error is deliberately toward keeping markets in. Two reasons, and the asymmetry is not the same one §4/§5 reasoned about:

  • Raising it buys nothing against the failure it exists for. A score computed from no data is fully prevented at $5. Every dollar above that excludes markets that genuinely can be assessed, in exchange for nothing.
  • Excluding a market is a much stronger action than excluding a reserve. A filtered reserve leaves a scored market and a published disclosure naming what was suppressed. An excluded market has no entry at all — no score, no factors, no disclosure, nothing for a reader to disagree with. §4/§5 already call hiding a small-but-real reserve "strictly worse" than leaving a dust one in; at market scale that error hides everything at once.

What this floor guarantees — and what it does not. It guarantees only that a published number was computed from something rather than nothing. It is emphatically not a quality bar: a market holding $50 clears it, and its score would still be close to meaningless. Saying so plainly matters more than the threshold does, because a floor that sounds like a meaningfulness test while being a scorability test is worse than no floor.

⚠️ A separate question this deliberately does NOT answer: is a scorable market worth listing? K2's primary market is a registered, ranked entry holding $1,781. It clears this floor by three orders of magnitude and is still small enough that a reasonable person could ask whether ranking it beside a $185M pool conveys what the ranking appears to convey.

That is a curation question — what belongs in the registry — not a question about whether a number can be computed, and answering it with a threshold in this document would dress an editorial judgment as a measurement. It also has a consequence a scoring threshold does not: any such bar set above $1,781 would delist a live entry, breaking a public URL (/protocol/kinetic) and orphaning a published history. Flagged here, resolved nowhere yet.

⚠️ An excluded market has NO score. It does not have a score of zero. This is the market-level form of the warning under §4/§5 about "cannot assess" quietly becoming a number again, and it is the whole reason the floor is written down.

  • It must mean: the market is not registered. If a registered market later falls below the floor, its published safetyScore must become null — the never-scored representation the API already defines and the dashboard already renders as an em dash.
  • It must never mean: a score of 0 (which renders in the danger band and says the opposite of what is true), a score of 100, or — the live hazard — registering the market and letting the five factors fall to their can't-assess branches, which is exactly what the shipped code does today and exactly how an empty market publishes a 0.

Enforcement, stated honestly: this floor is currently enforced only by the decision not to register such a market. No code path implements it. Nothing in an adapter, the indexer or the store can express "this market is not scorable" as distinct from "this market scored 0", so a registered market that drained below the floor would keep publishing a number today. Closing that needs a distinct not-scorable outcome through Adapter and RunRecord; it is filed in ROADMAP.md rather than implied to exist here.

This does not bump lending's methodology version. It moves no published number: every market Stenion currently scores — both Blend pools and K2's primary market — clears the floor, so no stored score is computed differently and none becomes non-comparable. It documents a precondition on what gets scored at all, which is additive.


5. utilizationSafety — headroom below the configured cap (weight 0.20)

What it measures: how close live utilization is to the protocol's own on-chain utilization stress line — the point the protocol itself defines as "borrowing should stop growing here." Approaching it is a concrete, protocol-defined stress signal.

Formula — headroom below the protocol's utilization line, worst reserve:

For each reserve with supplied > 0 and cap > 0 that passes the minimum-size filter:
    util = borrowed / supplied            # computed LIVE from balances, not a config field
    headroom = clamp( (cap − util) / cap × 100 , 0, 100 )

utilizationSafety = min(headroom) across all such reserves    # worst reserve wins

The minimum-size filter is §4's, unchanged and applied identically here — one rule, both factors.

Edge cases, all → 0, for the same reason as §4: no reserve with supplied > 0, every reserve excluded by the minimum-size filter, and no reserve with cap > 0. The second is the sharper one — reserves can hold real debt while declaring no utilization ceiling at all, and grading that as full headroom would measure distance to a line nobody set. The two are reported with different detail strings, since "the pool is empty" and "the pool declares no ceiling" are different findings.

cap is per-protocol — it is always the protocol's own on-chain utilization parameter, never a Stenion constant. Which parameter that resolves to:

Protocolcap sourceMeaning of the line
Blendper-reserve max_util (ResData/ResConfig, 7-dec fixed point → max_util / SCALAR_7)a hard throttle — Blend throttles and eventually pauses borrowing as utilization nears max_util
Kinetic (K2)OPTIMAL_UTILIZATION_RATE = 0.80 (contracts/shared/src/constants.rs)the interest-rate kink — past 80% util, K2's Aave-V3 rate curve steepens sharply to discourage further borrowing

Why this anchor (the strongest in the set): the threshold is not a Stenion constant at all — it is the protocol's own on-chain parameter. The formula grades each reserve against the exact line the protocol configured, so the "danger line" is set by the protocol, not by us. This is the pattern every continuous factor should aspire to. Worst-reserve selection is deliberate for the same reason as liquidity — the binding constraint is the single reserve closest to its line.

⚠️ Two honest caveats on the K2 anchor (flagged, not hidden):

  1. K2's kink is a rate inflection, not a hard pause — past 80% util K2 keeps lending (just expensively), whereas Blend's max_util is an actual throttle. The two lines mean slightly different things; the formula treats "distance to the protocol's declared utilization ceiling" uniformly, which is the intended abstraction.
  2. OPTIMAL_UTILIZATION_RATE is read as K2's global default (0.80). Per-reserve kink overrides, if any, live in K2's interest_rate strategy contract, which is out of scope in the audited source (code-423n4/2026-04-k2) and so not independently verifiable — if a reserve overrides the default this factor uses the documented 80%, not that reserve's exact kink. Revisit if K2 exposes a readable per-reserve optimal-util.

Dex

Everything from here to the end of this file is the DEX rulebook and the DEX rulebook alone — its version changelog, its two factors, what it refuses to score, and what it defers. Nothing in index.md is lending-specific and all of it applies here; nothing in lending.md may be assumed to hold for this category, and nothing here may be assumed to hold for lending.

Which markets are scored under it: one — Aquarius's XLM/USDC constant-product pool (CA6PUJLB…, registry id aquarius-xlm-usdc). This section was published and reviewable before any market was scored under it, which ../TAXONOMY.md says is the point of writing it first. It was admitted as a gate-checked submission against Aquarius on Stellar mainnet in two reviews — the factor set first, the weight table second — then implemented in adapters/aquarius/ and registered.

One market, and the reason is the indexer rather than the rulebook. Aquarius runs 340 pools across 304 token sets — read from the router's own get_pools_for_tokens_range at ledger 64,182,824 on 2026-08-29 — and every one of them is scorable under the rules below, because neither factor here is size-sensitive (see Size floor: none, and none pending). What limits the registry is that one scoring cycle runs inside a 60-second serverless ceiling and fits five markets, of which four were already lending. The other 339 are published as assessed and unregistered on the registry, under coverage.ts's awaiting-capacity status — never as a score, and never as a finding about them.

This rulebook is complete: two factors, each with a formula, and a reviewed weight table. The table was deliberately absent when the category was admitted and landed in a review of its own — a weight can only ever be a type-(b) unvalidated judgment call, and arguing one alongside the argument for the factor set would have buried it. Both halves are here now, in Factor weights and The two dex factors, and CATEGORY_FACTORS.dex in ../core/src/weights.ts carries status: 'published' to match. Every Stenion-chosen number in either formula is listed in Unvalidated judgment calls.

The category ships with TWO factors, not the three the submission proposed. depthSafety is deferred by question A, which is resolved rather than open. Gate 0 is re-argued for the two that remain in Gate 0, re-argued.

There is no size floor, and none is pending. Neither surviving factor is size-sensitive, so there is nothing for a floor to protect — see Size floor: none, and none pending. This is a consequence of question A's resolution, not an open item.

Everything below was read from mainnet, not from documentation. That distinction is load-bearing for this category in a way it was not for lending: github.com/AquaToken/soroban-amm — the repository Aquarius's own audit scope links to — returns 404 as of 2026-08-27. There is no source to read. Every interface claim here comes from the contract spec in the deployed wasm (getContractMethods) and from instance storage, which is what Gate 8 asks for anyway.

Two dated read sets appear below, and each reading says which it is from. The census readings — the 340-pool survey, the role structure, the issuer-flag sample, the kill-switch state — are mainnet ledger 64,152,946, 2026-08-27T20:00Z, taken when the category was admitted. The worked example is computed from the mainnet fixtures captured on 2026-08-29T09:35Z, which live in adapters/fixtures/aquarius/ and are checked into the repo, so its arithmetic can be re-derived rather than taken on trust.

Version changelog

Every scored run is stamped with the rulebook version that produced it (risk_scores.methodology_version, from the dex entry in METHODOLOGY_VERSIONS in ../core/src/category.ts), and it is surfaced on the API's protocol detail and on each history point. Versions are per category and counters are independent, so this changelog is dex's. dex v1 and lending v1 are not two editions of one rulebook and one is not older than the other; they are two different rulebooks that each start counting at 1. The category is stored beside the integer because the integer alone does not identify a rulebook.

VersionEffectiveChange
12026-08-29The initial two-factor model as documented here — adminKeySafety (role posture and the upgrade reaction window) and assetControlSafety (issuer freeze and clawback) — weighted 0.55 / 0.45. depthSafety is deferred (question A, option 4) and no size floor is needed by either factor that ships. The factor set was admitted first and the weight table reviewed second; both are version 1, because no score was ever published under the factor set alone. Live since aquarius-xlm-usdc was registered on 2026-08-29; every stored dex row carries it.

Changed under v1, without a bump: 2026-08-30, a rate limit is no longer a reading. Aquarius was capturing a Horizon read that never reached an answer — an exhausted 429, a dropped connection, a 5xx — as a localized cannot-assess and scoring it 0. It is now a failed run. Recorded here rather than as a version 2 because What bumps the version says a fix that makes the implementation match the documented rule does not bump: the affected scores were wrong under this rulebook, not produced by a different one. No formula, tier or weight moved. Full reasoning in A rate limit is not a reading.

One row, and it describes the rulebook every stored dex score was produced under. It was published before any of them existed, which is deliberate: Gate 7 requires a category to arrive at version 1 rather than acquire a version once it starts producing numbers, so the counter exists from the moment the rulebook is published. The first stored dex row carried version 1 — settling the weight table did not bump it, because there was no prior published number for a weight to make incomparable: a rulebook that could not compute a score cannot have produced one that a weight made non-comparable. From that first stored row onward the ordinary rule in What bumps the version applies without exception, and the next change to either weight is a version 2.

History is not backfilled across a bump, and cannot be, here as everywhere: risk_scores stores only outputs, never the raw on-chain inputs a run was computed from.

Factor weights

Dex's weights, and dex's only. This table is the published face of CATEGORY_FACTORS.dex in ../core/src/weights.ts, which is where the adapter reads them from — no adapter contains a weight of its own, and core/src/scoring.test.ts parses this table and fails if the two disagree in either direction.

FactorWeight
adminKeySafety0.55
assetControlSafety0.45
Total1.00

Worked example — the Aquarius XLM/AQUA constant-product pool CCSY43EHJAHT3NQDYKAMJXRFBEEH7OXDL3J3VNGO33UUSEXWNN27GBIZ, from the mainnet fixture captured 2026-08-29T09:35:43Z (adapters/fixtures/aquarius/constant-product-mainnet.ts), dex methodology v1:

10×0.55 + 70×0.45 = 37.0 → 37

Both terms, derived from that fixture's own fields so the arithmetic is checkable rather than asserted:

adminKeySafety = 10. Role posture is the minimum over the seven roles (§1):

RoleRead from the fixtureBaseActivity penaltyScore
AdminsignerCount 3, high_threshold 290−30 (96 ops)60
EmergencyAdminsignerCount 1, high_threshold 040−0 (0 ops)40
EmergencyPauseAdminsignerCount 1, high_threshold 040−0 (0 ops)40
PauseAdminsignerCount 1, high_threshold 040−0 (0 ops)40
OperationsAdminsignerCount 1, high_threshold 040−0 (0 ops)40
RewardsAdminsignerCount 1, high_threshold 040−30 (200 ops)10
SystemFeeAdminsignerCount 1, high_threshold 040−30 (200 ops)10

rolePosture = min(…) = 10. The fixture's upgrade.deadline is 0n on both the pool and the router, so upgradeCeiling = 100 and the ceiling does not bind: min(10, 100) = 10.

assetControlSafety = 70. The minimum over the pool's two reserve tokens (§2):

Reserve tokenRead from the fixtureScore
XLM CAS3J7GY…SAC, issuer.status = noIssuer — no issuer account exists100
AQUA CAUIKL3I…SAC, issuer GBNZILST…, all four flags false, not authImmutable70

min(100, 70) = 70.

The other three fixtures, for the spread — same weights, same formulas, computed the same way, from the same 2026-08-29 capture. All four pools read the same admin posture, so assetControlSafety is the only term that moves:

FixturePooladminKeySafetyassetControlSafetyScore
constant-product-mainnetXLM/AQUA107037
concentrated-mainnetXLM/AQUA107037
stable-mainnetUSDC/USDx/yUSDC1040 (Circle USDC is auth_revocable)24
wasm-token-mainnetUSDC (SAC) / USDC (wasm)1040 (the wasm token is excluded, route (a))24

The weights are an unvalidated judgment call, not an external fact. adminKeySafety carries more because its subject is the pool's own code and role set: a compromise there reaches every token the pool holds, the fee it charges, whether it trades at all, and the code path an LP's withdrawal runs through. assetControlSafety reaches only the balances of one issuer's own asset and cannot touch the pool's code — a pool holding one flagged token and one clean one has a fraction of its value exposed, and the AMM's withdraw path still works.

It is close behind rather than a minor term because it is the one failure Aquarius cannot mitigate and the LP gets no warning for: a clawback is a single issuer transaction with no window at all, while a code change is announced by UpgradeDeadline before it lands. Those two arguments nearly cancel, which is why the gap is 0.10 rather than lending's spread — the ordering is the claim, and the gap is deliberately small.

Direction of error: if the split is wrong it is wrong by under-weighting issuer control. A reader who holds that unmitigable-and-unannounced should dominate would push toward 0.45 / 0.55, and there is no external framework that anchors against them. Full label and the rest of the list in Unvalidated judgment calls.

They were not chosen to make the scores spread out, and must not be. All 340 pools shared one admin posture in the 2026-08-27 census, and the four pools captured on 2026-08-29 still did — so adminKeySafety discriminates between Aquarius markets not at all today and assetControlSafety carries the whole variance, as the four-fixture table above shows. Down-weighting the constant factor to widen the registry's range would be calibrating a category rulebook against one protocol's current data, which is the failure this document refuses everywhere else. A weight states which failure matters more, not which reading varies more.

Both steps of the two-step admission are now done. ../TAXONOMY.md admits a category in two reviews, and this one passed through both:

Step 1 — the factor setWhich failures are scored, what each factor's on-chain anchor is, what was rejected. Reviewable on its own, and reviewed.
Step 2 — the weight tableWhat each factor's share is and what each formula computes, argued as judgment calls and labelled as such. Done.

Between the two, CATEGORY_FACTORS.dex declared status: 'pendingWeights' and its factor entries carried no weight property at all — not a zero, not a placeholder — so reading a weight off one was a compile error rather than an undefined that would become NaN inside scoreFactors and publish a confident-looking 0. It now declares status: 'published', and the table above is pinned against it by core/src/scoring.test.ts in both directions: a weight edited here alone, or there alone, fails.

The two dex factors

For each factor: the exact raw on-chain data that feeds it, what it detects, and why it is here. Every anchor names a contract and a method or storage field, per Gate 8.

Two, not five, and the missing three are not an omission. Three of lending's five have no referent in a spot AMM at all:

Lending factorStatus for a spot AMMWhy
utilizationSafetyNo referentIt grades distance from a protocol-declared borrow cap. An AMM has no borrow ledger and no cap; nothing resembling one appears in any of the three pool wasms.
liquiditySafetyNo referent(supplied − borrowed) / supplied is identically 1 for every AMM pool. The formula would publish 100 for all 340 pools, from no information.
oracleSafetyNo referentAquarius reads no price feed anywhere — established exhaustively below, not by failing to find one. Price comes from reserves, or from tick state.
collateralSafetyMisleading if reusedHHI over reserves. A two-token pool's reserve split is its price, so the HHI is dominated by the two assets' unit prices rather than by any concentration risk — it would grade a pool worse for pairing assets of unequal price.
adminKeySafetyApplies, same key"Who can change the rules" is genuinely the same question. Kept under the same name, computed from different data — see below.

Aquarius reads no oracle, and this was established rather than assumed. The exported function list of the router, all three pool wasms, the plane, the liquidity calculator, the config storage and the reward-boost feed were each read out of the deployed wasm. None contains a price read, and no price-feed address appears in any instance storage. The one contract whose name suggests otherwise — RewardBoostFeed CBKCROE56TU2FTT3C5CVN676PYVLTOQUQDHHH57GLWDY5VOKSCZPGOFN — exports exactly total_supply() and set_total_supply(operations_admin, …) and holds TotalSupply = 642372689091226311. It is a locked-AQUA supply feed for reward boosting; it is not a price oracle. So oracleSafety is not "ungradable" in Gate 2's sense — it has nothing to be ungradable about, which is why it is absent rather than disclosed.

Gate 0, re-argued for two factors

The submission proposed three factors and this rulebook ships two: adminKeySafety and assetControlSafety. depthSafety is deferred — see question A. Question A's option 4 said plainly that "a two-factor category is thin enough that Gate 0 should be re-argued before accepting it," so it is re-argued here rather than assumed to carry over.

Gate 0 asks whether the category's factors answer a question no existing category's factors already answer. Each of the two does, on its own:

  • adminKeySafety — the two-step upgrade reaction window. Aquarius's upgrades are commit_upgrade → wait → apply_upgrade, and while a deadline is pending UpgradeDeadline − now is exactly how long an LP has to withdraw before the code under their money changes. Lending's adminKeySafety cannot see this: it reads one admin account's signer set and threshold, which says who could act and says nothing about how much warning anyone gets. The factor shares lending's key because it is the same question — but a lending market's answer to it is computed from a different quantity, and the reaction window is a failure mode lending's version has no access to.
  • assetControlSafety — issuer-level freeze and clawback. A pool can be created permissionlessly against any token; if that token is a SAC whose issuer has auth_revocable or auth_clawback_enabled set, the issuer can freeze or seize the pool's balance and an LP's exit stops depending on the AMM's code at all. No lending factor detects this, and it is not a rephrasing of one: collateralSafety measures concentration among assets, not whether a third party outside the protocol can take them.

Two is enough to clear the gate, because Gate 0 is a per-factor test and not a headcount. It fails a factor that duplicates an existing one; it does not set a minimum. Both survivors detect failures lending's five cannot see, from readings with real live variance — Circle's USDC issuer has auth_revocable: true, a second asset also called USDC sits in the same registry, the owner key is a 2-of-3 multisig while six other roles are lone keys, and none of that is visible from any lending factor.

What this does NOT claim, stated because the omission is the part that could mislead. Leaving depthSafety out is not a judgment that execution cost is unimportant to DEX risk — it is close to the most important thing about an AMM, and the category's own Gate 0 sentence asks whether a trader can get out at the size they are actually trading. This rulebook currently answers the second half of that sentence (can an LP's capital leave on terms the pool's own code decides) and not the first. The reason is narrow and it is about anchoring, not importance: Aquarius publishes no unit of value, so any trade size we simulated at would be a number Stenion chose rather than one the chain states, and Gate 1 exists to keep exactly that out of a published score.

The consequence is that a dex score is a governance-and-asset-control score, not a liquidity score, and it must not be read as evidence that a pool is deep. Two things follow, and both are obligations rather than caveats:

  • Depth is still published, as a route-(a) value: null disclosure carrying the estimate_swap readings — so a reader gets the measurement without it being graded. A deferred factor is not a hidden one.
  • The scores will cluster, and that is a property of the data rather than a defect. All seven roles read identically across the router and all 340 pools on 2026-08-27, so adminKeySafety does not currently discriminate between Aquarius pools at all; assetControlSafety is the only factor that varies, on the tokens a pool holds. Two pools with the same tokens will publish the same number. Anyone registering more than a handful of Aquarius markets should read that as the registry reporting the truth — these pools really do share one admin posture — and not as a ranking.

Deferred, and not declared: depthSafety

depthSafety is not one of this category's factors. It was proposed as the third and is deferred by the resolution of question A. It is not declared in CATEGORY_FACTORS.dex, is not weighted, and is not computed: a declared factor with no formula is a promise the rulebook does not keep, and weights.test.ts asserts the key is absent so it cannot be added back without also publishing the rules for it.

What follows is kept, in full, because the deferral is about one missing input and not about the measurement being wrong. Everything here holds the day a unit of value exists; discarding it would mean re-deriving it later from a repository that no longer exists.

Anchors, when it lands: estimate_swap(in_idx: u32, out_idx: u32, in_amount: u128) -> u128 on the pool contract, with get_reserves() and get_fee_fraction() on the same contract.

The failure it would detect: a trader cannot get out at the size they are actually trading. Nothing in lending's five measures execution cost, because a lending market has no execution — a withdrawal is at par or it is refused. That failure is real and this category does not currently measure it — see Gate 0, re-argued for what is and is not being claimed by leaving it out.

estimate_swap is a pure simulation call that runs the pool's own curve — constant product for standard, Curve-style stableswap for stable, the tick walk for concentrated. So the slippage figure is computed by the contract being scored rather than modelled by us, which is the strongest form Gate 8 admits: there is no Stenion-side curve implementation to drift from the deployed one. Both trade directions are simulated and the worse direction is the one that counts, on the same convention every lending factor uses for reserves — the binding constraint is what a reader needs.

The fee is inside this number already, via get_fee_fraction(), because a round trip pays it. That is the only place fee belongs — see Fee tier as a factor below.

Live readings on the XLM/USDC pair, 2026-08-27, showing the discrimination the factor exists to produce — two pools on the same pair, one materially deeper, both figures produced by their own contracts:

PoolType, fee1,000 XLM in1,000,000 XLM inImpact
CA6PUJLBYK…standard, 10 bps0.186277 USDC/XLM0.168205−9.70%
CBBMQBNHB2…concentrated, 10 bps0.1862650.175584−5.74%

Cannot-assess would resolve to 0, the unsafe end. A pool whose estimate_swap reverts, or whose reserve is zero, would score 0 — never 100, and never a skipped factor, matching collateralSafety's existing treatment of an empty set.

What is NOT settled, and is why this is deferred: the trade size it simulates. Everything above describes how the cost is measured, and all of it holds. At what size has no answer in Aquarius's own terms — a pool-relative size makes the factor degenerate across 272 of 340 pools, and an absolute size needs a unit of value Aquarius does not have. See question A. No adapter may implement this factor until that changes.

1. adminKeySafety — seven roles, and how long you get to react to a code change (weight 0.55)

Anchors: get_privileged_addrs() -> Map on the router and on every pool; the UpgradeDeadline and FutureWASM entries in contract instance storage; Horizon /accounts/{G…} for each role's signer count, thresholds and recent activity.

This is the same factor key lending uses, and that is a decision, not an oversight. It is recorded here so it is not re-litigated, and it was an open question when the category was admitted:

Gate 0 rejects a factor that is "an existing factor rephrased, rescaled, or renamed for a new audience." "Who can change the rules under my money" is not a rephrasing of lending's question — it is the same question, and the honest way to say so is to use the same name. The data underneath is entirely different: lending reads one admin account's signer set and threshold, while this reads seven named roles plus a two-step upgrade deadline. A separate key — roleControlSafety, say — would have published two names for one failure and started the taxonomy fragmenting into synonyms, which is the drift the shared RiskFactorType vocabulary exists to prevent. So: one key, two computations, declared per category in CATEGORY_FACTORS.

A shared key is not a comparability claim. A dex adminKeySafety of 60 and a lending adminKeySafety of 60 were produced by different rules from different data and mean different things — exactly as the two categories' overall scores do. See Comparability.

The failure it detects that lending's version cannot: the upgrade reaction window. Aquarius upgrades are two-step. commit_upgrade(admin, new_wasm_hash, …) writes UpgradeDeadline; apply_upgrade(admin) refuses until that deadline passes; revert_upgrade(admin) cancels. When a deadline is pending, UpgradeDeadline − now is exactly how long an LP has to withdraw before the code under their money changes. It is anchored with no Stenion constant in it at all — the chain states the deadline and the chain states the time.

Read on 2026-08-27: UpgradeDeadline = 0 and FutureWASM == running wasm on the router and all 340 pools. No code change is scheduled anywhere.

The role structure, read on 2026-08-27 — identical across the router and all 340 pools:

RoleAccountSignershigh_thresholdops in 30d
Admin (owner / upgrade)GAV5FBMKD2ZF4X2MGWDNQYUP7KFL7MRM6HZBY7HKQLB4BRHSCCX5J6VS3296
EmergencyAdminGCGZ6E5RBUKLNB4VZ5RC65C4QMBSBJ3COVRRJCWAMCXJC36LB7YYWEKM100
EmergencyPauseAdminGA6MVTGQDCJPP27IAMG6PSDTWOJYTD3NUTLR2W54ADBCBY7OID5YUDSI100
PauseAdminGA6MA665XVKHTQUZVSMUKUPGT7OREJNCLAZ5ZEH5CXPKYTWFJKZ3YSEK100
OperationsAdminGBVQPX2LQ55HLRMLIWBEYVVQL3SZ5RFPRKYLRSLZU4XRIWAXW2KQIMMD100
RewardsAdminGCXYKA3BM574WC6TWESEDUGUJTNQ5SVCFMHWLQ634H5FTE7FYPV3JH3X10200
SystemFeeAdminGB57YDVGLL2BAVOXHPXYCZR77J4MLPLMGJKFMTUKMHFI2AEGS4SGGW7N10200

The owner key is a 2-of-3 multisig; the other six are lone keys. That is a real, checkable posture and it is not the one Aquarius's own auditor recommended — Certora's Appendix A asks for the Pause Admin to be "a multisig or DAO" and for the owner key to be kept offline. Both halves are readable, and both disagree with the recommendation. Publishing that disagreement is the factor doing its job.

Formula — a per-account tier, minimised over the roles, then capped by the upgrade state.

Two components, each computed from the readings above, combined by taking the worse of the two.

Component 1 — role posture, from get_privileged_addrs() and Horizon /accounts/{G…}:

base = 90   signerCount > 1 AND high_threshold > 1        # N-of-M multisig
       40   a classic account that is not a multisig      # single master key
        0   a `C…` contract address, or the lookup failed # cannot assess -> unsafe end

activityPenalty = min(30, recentOps × 3)     # recentOps = operations in the last 30 days
accountScore    = clamp(base − activityPenalty, 0, 100)

roleScore   = min(accountScore) over the addresses that role holds
rolePosture = min(roleScore) over the seven roles

Component 2 — the upgrade reaction window, from UpgradeDeadline and the fetch timestamp:

upgradeCeiling = 100   UpgradeDeadline == 0                # no code change is scheduled
                  40   UpgradeDeadline >  fetchedAt        # scheduled; the window is still open
                   0   UpgradeDeadline <= fetchedAt, non-zero
                                                           # matured: applicable at the next
                                                           # ledger, with no warning left
adminKeySafety = min(rolePosture, upgradeCeiling)

Why these numbers:

  • The per-account tier is lending's, adopted rather than re-argued — 90 for an N-of-M multisig, 40 for a single master key, −3 per operation capped at −30, exactly as lending §3 publishes them. The reading underneath is identical on both sides: Stellar's account threshold model and a 30-day operation count, from the same two Horizon calls. Inventing a second set of integers for the same reading would be two rulebooks for one question — which is the drift the shared adminKeySafety key exists to prevent, applied to the numbers rather than to the name. So the split is anchored exactly as it is there (a 1-of-1 key is a single point of unilateral compromise; signerCount > 1 AND high_threshold > 1 provably is not, per Stellar's own threshold model), and the exact integers remain a labelled judgment call exactly as they are there.
  • The contract-address branch is 0 here and 60 in lending, and that divergence is the one deliberate departure. Argued in Cannot-assess below: a contract admin is a known, named structure lending chose not to grade, while an Aquarius role that is not a classic account is a structure we did not expect and cannot describe.
  • The 40 in upgradeCeiling is not a new constant — it is the single-master-key tier, reused. While a code change is scheduled, the LP's only remaining protection is a countdown whose length cannot be read at all, so the market is graded no better than one whose rules a lone key can change. The rule introduces no integer this document did not already publish.
  • The 0 once the deadline has matured is a reading, not a choice. What this component grades is UpgradeDeadline − now, and once that is non-positive the measured warning is zero. The chain states the deadline and the chain states the time; there is no Stenion constant in that branch.
  • FutureWASM is read but does not enter the formula. Every contract read on 2026-08-29 carried a FutureWASM equal to its own running hash, so presence is the quiescent state rather than the signal — UpgradeDeadline is what says a change is scheduled. FutureWASM and whether it differs from the running hash are published in the factor's detail string, because a reader wants to know which code is staged, and they are not graded because the deadline already carries the fact.

Two of the choices above are unvalidated judgment calls — combining the seven roles by min, and the pending-upgrade ceiling. Both are labelled, with their direction of error, in Unvalidated judgment calls.

Two limits, stated rather than papered over:

  • The timelock duration is not readable, and that is why upgradeCeiling is a state rather than a curve. ADMIN_ACTIONS_DELAY is a compile-time constant with no getter, and the source repository is gone. The chain states the deadline and the chain states the time, so the factor can say whether a window is open, say when one has matured, and publish the seconds remaining — what it cannot do is say what fraction of the window is left, because it never learns the whole. Any grading of the remaining window's length would need a threshold in seconds that Aquarius does not state, which is an invented number and ground rule 4 forbids it. So the duration itself takes Gate 2 route (a): a value: null disclosure component saying it is not readable from the contract, published beside the factor rather than folded into it.
  • The Emergency Admin can bypass the delay. In Aquarius's own words, in their response to Certora's H-01: "In the case of system vulnerability fixes, delay may be bypassed by the Emergency Admin role." So the reaction window is conditional on one key choosing not to skip it — and that key is one of the six single-signer accounts above. Also route (a): disclosed beside the window, never silently folded into it, because "a window exists" and "the window is unconditional" are different claims and only the first is true.

Why route (a) and not the other two, for both, as Gate 2 requires be argued rather than asserted. Route (b) — a precondition that leaves the protocol unscored — would refuse to score every Aquarius market over two facts that are the same for all 340 of them and that neither changes nor discriminates: nothing would ever be published, and a reader would learn less, not more. Route (c) — live ungraded state — is for readings that change between cycles, and neither does: ADMIN_ACTIONS_DELAY is a compile-time constant in code we cannot read, and the Emergency Admin's bypass is a property of the deployed contract's authorisation logic. Both are fixed facts about an unreadable quantity, which is exactly the shape route (a) exists for.

2. assetControlSafety — can a third party freeze or seize what the pool holds (weight 0.45)

Anchors: each reserve token contract's instance executable (a Stellar Asset Contract's is contractExecutableStellarAsset) and its METADATA instance entry, whose name is CODE:ISSUER for a classic asset and the bare string native for XLM, then Horizon /accounts/{issuer} for that issuer's flags.

Corrected from the submission, which named an AssetInfo storage key. There is no such key; The issuer was found in METADATA by probing the deployed token contracts. The Gate 8 claim is unchanged — the issuer is read from the token contract's own instance storage, not from an aggregator — only the key name was wrong. The native counter-example gets sharper as a result: XLM's METADATA.name is the bare string native, which is why a SAC is detected from its executable and never from the shape of its name.

The failure it detects, and it has no analogue in lending's five: an Aquarius pool can be created permissionlessly against any token. If that token is a Stellar Asset Contract whose issuer has auth_revocable or auth_clawback_enabled set, the issuer can freeze or seize the pool's own balance — and an LP's exit stops depending on Aquarius's code at all. No amount of depth and no admin posture protects against it. Aquarius's auditor named this too: Certora M-02, "Lack of scam protection for AMM Users."

Readable, with real variance. Of 205 distinct tokens across the 340 pools, 196 are Stellar Asset Contracts (executable = contractExecutableStellarAsset, issuer parsed out of METADATA) and 9 are wasm contracts. Sampled issuer flags, 2026-08-27:

  • USDC:GA5ZSEJY… (Circle, home_domain = circle.com) — auth_revocable: true. Same for EURC:GDHU6WRG….
  • LSP:GAB7STHV… — auth_immutable: true, the strongest reading available: the flags can never change.
  • AQUA, yXLM, USDx, EURx, BLND, SSLX, WHLAQUA, XRF — all four flags false.
  • And a second asset also called USDC, issuer GCBYVQH3…, home_domain = mirrasets.com, sitting in the same registry as Circle's. That is Certora M-02 in the live data, not in theory.

Formula — a per-token tier on the issuer's flags, minimised over the pool's reserves:

Per reserve token:
    100   a SAC with no issuer account at all — native XLM
    100   auth_immutable set, with auth_revocable AND auth_clawback_enabled both clear
     70   auth_revocable and auth_clawback_enabled both clear, auth_immutable not set
     40   auth_revocable set, auth_clawback_enabled clear
      0   auth_clawback_enabled set
      0   the token IS a SAC and the Horizon read for its issuer failed
      -   not a SAC: excluded from the computation, route-(a) disclosure (see below)

assetControlSafety = min(tokenScore) over the graded tokens
                     0 when no token is gradable

The order the tiers are tested in is load-bearing. Clawback is tested before revocable, and both before immutable, because auth_immutable freezes whatever the flags currently are — it is a credit only when what it freezes is clean. An issuer that is both immutable and revocable has made its freeze power permanent and scores 40, not 100. Testing immutability first would invert that.

auth_required moves no number, deliberately. It gates who may acquire the asset; it does not let an issuer touch a balance that already exists. The power to freeze one is auth_revocable and the power to take it is auth_clawback_enabled. auth_required is read and published in the factor's detail string, because it describes the asset a reader is looking at, and it is not graded because it does not answer this factor's question.

Why these numbers. The ordering is anchored to what Stellar's account flags let an issuer do, which is a protocol-level fact and not a Stenion preference: seizure (auth_clawback_enabled) strictly dominates freezing (auth_revocable), which strictly dominates an issuer holding neither, which is in turn a weaker statement than an issuer that can never acquire either — and weakest of all against an asset with no issuer account for anyone to act from. The spacing — the 40 and the 70 — is an unvalidated judgment call, labelled in Unvalidated judgment calls.

Why 70 and not 100 for an issuer whose flags are clean. Because auth_revocable is one SET_OPTIONS away for any issuer that has not set auth_immutable, so "clean today" is a strictly weaker statement than "cannot become dirty". That distinction is the whole reason the second asset also called USDC (GCBYVQH3…, home_domain = mirrasets.com) is worth publishing about: its flags read exactly like a well-run issuer's, and in both cases the reading is one transaction from changing. Scoring it 100 would publish the reading as a guarantee.

The 9 wasm tokens have no issuer-flag equivalent (SolvBTC, xSolvBTC, BnUSD, XAUM, and wasm USDC/USDT variants). They take Gate 2 route (a): a value: null disclosure component naming the token and saying the read does not apply. Not a silent pass, and not a silent 0.

Why not the other two routes, as Gate 2 requires be argued rather than asserted: route (b) — leaving the protocol unscored — would drop every pool containing any of nine tokens over a signal that is absent rather than broken, which is far stronger than the finding warrants. Route (c) — live ungraded state — is for state that changes, and "this token is not a SAC" is a fixed property of the contract's executable, not a reading that can flip between cycles.

Cannot-assess, stated per factor — and what is a failed run instead

Gate 2 requires every "cannot assess" branch to resolve to the unsafe end of the scale. That is one sentence in the checklist, and it needs an answer per factor rather than one example standing in for the set — so both factors state their own below, and neither inherits the other's by implication.

The rule, stated once and applied to both: a partial or unreadable input scores 0, never a skip and never a pass. Where a read leaves any part of a factor's input unavailable — a call that reverts, a response shorter than expected, a lookup that fails for one subject — the factor resolves to 0, the unsafe end. It is never omitted from the factor map, never defaulted to a neutral middle, and never allowed to score well because there was less to check. This is ground rule 4, and lending's 2026-08-16 correction is the precedent: two factors there returned 100 from an empty filtered set, which published "maximally safe" from no data, and both were changed to return 0.

One boundary, and it is the difference between publishing a 0 and publishing nothing. A whole-endpoint outage — Soroban RPC unreachable, Horizon down, nothing decodes — makes the adapter throw, and the indexer records a failed run and publishes no score for that cycle (CLAUDE.md's error-handling rule; error handling lives in the indexer, never per adapter). That is not a cannot-assess branch and must not be graded 0, or a blip in our own network path would publish "dangerous admin control" across every pool at once. Everything localized — one call, one role, one issuer — is a cannot-assess branch and takes the 0.

"Localized" is not a count of failed reads; it is a claim about what the failure is a statement about. The two words that separate the branches are the subject answered. A read whose subject answered — an issuer account Horizon says is not there, a contract that reverted — produced a fact about the pool, and a fact about the pool is what a factor is for. A read that never reached an answer — the endpoint refused us for rate, the connection dropped, Horizon returned a 5xx about itself — produced a fact about Stenion's read path on that cycle, and no such fact may move a protocol's number. Both used to be called "the lookup failed"; they are opposite claims, and the one that is not about the protocol fails the run. The argument, the options weighed against it, and what it does and does not change are in A rate limit is not a reading below.

FactorCannot-assess branchResolves to
adminKeySafetyget_privileged_addrs() reverts; returns fewer than the seven roles; a role is ungradable; the contract carries no UpgradeDeadline entry at all0
assetControlSafetyHorizon answers about a SAC issuer and the answer is not a gradable account; no reserve token is readable at all0
(either factor)The endpoint never answered: a rate limit we could not wait out, a dropped connection, a 5xxNot a branch — a failed run, and no score published for the cycle
(disclosures)Timelock duration; Emergency Admin bypass; the 9 non-SAC wasm tokensNot a branch — route-(a) value: null, moves the number neither way

adminKeySafety. Three ways the input can be incomplete, all resolving to 0:

  • get_privileged_addrs() reverts on the router or on the pool. The whole role structure is the factor's input, so a revert leaves nothing to grade. 0, not a skipped factor — "we could not read who controls this pool" is a statement about the pool, and the unsafe end is the only honest place for it.
  • It returns fewer than the seven expected roles. Seven were present across the router and all 340 pools on 2026-08-27 (Admin, EmergencyAdmin, EmergencyPauseAdmin, PauseAdmin, OperationsAdmin, RewardsAdmin, SystemFeeAdmin). A short map means either an unexpected contract version or a role we cannot see, and grading the roles that did come back would publish a posture assessment of an admin set we know is incomplete. 0.
  • A role reads but cannot be graded — its address is not a classic G… account, so there is no signer set, threshold or activity history behind it. 0.
  • The contract carries no UpgradeDeadline entry at all — the raw shape records this as deadline: null, which is a different statement from 0n and must not be collapsed into it: one says the contract answered "nothing pending", the other says this contract does not keep the field the upgrade component reads. Both the router and pools of all three types carried it on 2026-08-29, so its absence would mean an unexpected contract version, and half this factor's input would be missing. 0, on the same reasoning as a short role map.

That third case argues with an existing precedent, and the resolution is recorded rather than assumed. Lending's adminKeySafety sends its equivalent to a clearly-flagged neutral 60, on the reasoning that a contract-held admin is a different structure rather than evidence of danger. dex does not inherit that baseline: lending's 60 is anchored to a contract admin being a known, named structure whose properties we chose not to grade, whereas an Aquarius role that is not a classic account is a structure we did not expect and cannot describe. A neutral score for an unexpected reading is an invented number, which ground rule 4 forbids. All seven roles were classic accounts on 2026-08-27, so this branch is unexercised today.

assetControlSafety. Two cases, and the distinction between them is the whole point:

  • Horizon answers about a SAC issuer's /accounts/{issuer} and the answer is not a gradable account — a 404, most plainly. The token is a SAC, so its issuer flags are exactly the thing this factor grades, and Horizon has told us there is nothing there to grade. 0 for that token, which then binds the factor on the usual worst-reserve convention. This is not the same as the 9 non-SAC tokens and must not be routed like them: there, the read genuinely does not apply; here, it applies and failed. Treating a failed read as "does not apply" would silently upgrade an unknown into an exemption, which is the single most dangerous confusion available in this factor. It is equally not the same as a read that never reached an answer — see A rate limit is not a reading.
  • No reserve token is readable at all — every reserve is a wasm contract, so every token takes the route-(a) disclosure and nothing is left to grade. A minimum over an empty set is 0, not 100. Same shape as lending's 2026-08-16 correction, written down before the adapter exists rather than after it publishes a 100.

The 9 non-SAC wasm tokens remain what they were: a route-(a) value: null disclosure naming the token and saying the read does not apply. A disclosure is not a zero and not a pass — the token is excluded from the computation, not graded badly by it.

A rate limit is not a reading

The question. Aquarius is the first adapter that scores through a failed read instead of throwing on it, and the 429-backoff work made the consequence concrete: when Horizon refuses us for rate and the attempt's retry budget is spent, the refusal arrived at scoring as a 0 with a reason string, indistinguishable at the factor level from "this issuer's flags could not be read." Those are different claims. One is about the pool. The other is about Stenion's infrastructure at the moment of the read, and it was moving a published number.

The decision: a rate-limit exhaustion, and every other failure that never reached an answer, is a failed run — not a scored 0. Nothing about the two formulas, the tier values, the weights or the worst-reserve convention changes. What changes is which failures are allowed to reach them.

The rule, in the form the code applies it. A read is a reading only when its subject answered, and each transport is asked that question in its own vocabulary:

TransportAn answer about the subject — captured, scoresEverything else — fails the run
Horizonan HTTP status that is about the account: 404, and any other 4xxa 429, any 5xx, a rejected fetch, a body that did not parse
Soroban RPCa simulation error — the contract itself refusinga rate limit, a dropped connection, an SDK decode failure

The classification lives in ../adapters/read-failure.ts, not in each call site, and it is tested there and in adapters/aquarius/fetch.test.ts against a loopback Horizon answering 404, 503, 429 and nothing at all.

Why, in three facts rather than a preference.

  1. A 429 carries no information about the protocol. Every one of Aquarius's 340 pools is equally rate-limitable, on a schedule set by our own request train against a free shared endpoint. A number that moves with it is reporting Stenion's cycle, and the registry has no column for that.
  2. It is not localized, which is the word the boundary above turns on. The retry budget is per attempt and shared by every call the attempt makes (../core/src/rate-limit.ts): once it is spent, the next refused read gets no retries at all. So one exhaustion silences the reads after it, adminKeySafety and assetControlSafety both resolve to 0, and the pool publishes an overall 0 — "dangerous admin control", from an endpoint that was busy. That is precisely the outcome the whole-endpoint boundary exists to forbid, arriving one pool at a time instead of all at once.
  3. The honest publication already exists and costs nothing to reach. A throw runs the indexer's own target retry with a fresh budget, up to STENION_RETRY_ATTEMPTS times; only if every attempt is refused does the cycle record a failed run, flagged rateLimited and carrying the 429s in risk_scores.error. The previous score stands with its staleness advancing, which says the true thing: we did not learn anything about this protocol this cycle.

The options that were weighed, and why each was not taken.

OptionWhy not
Keep the scored 0. Gate 2 resolves uncertainty to the unsafe end and does not ask why.Gate 2 governs a cannot-assess branch, and a cannot-assess branch is a reading. The rule it appeals to is ground rule 4 — never publish a number no data supports — and a 0 sourced from our own request rate is that number, not an application of it.
Carry the prior cycle's sub-reading forward and retry next cycle.risk_scores stores outputs, never the raw inputs a run was computed from, so there is nothing to carry: this needs a new persistence layer for per-sub-value state, and a score assembled from two cycles' readings is a fourth publication route nobody has argued for. Flagged as a separate design question, not forced into this pass.
Keep the 0 and disclose it in detail, plus a dashboard treatment.It fixes the ambiguity for a reader who opens the factor and not for the ranking, which is where the damage is: the pool sits last in the dex block with a footnote. Disclosure is the right response to something we chose not to grade; this is something we never measured.
Throw on a 429 only, leaving the other never-answered failures scoring 0.A dropped connection and a Horizon 5xx are the same claim about the same thing and were scoring 0 too — "Horizon down" was publishing a risk finding, against the boundary this document already wrote. Fixing the named case and leaving its siblings would have made the rule a special case instead of a rule.

What this deliberately does NOT do. Aquarius still scores through every failure that is a statement about the pool: a reverting get_privileged_addrs(), a short role map, an ungradable role, an issuer account Horizon says is not there. The score-through-failure design is intact — it was never a design to score through our own failures. Soroban RPC's 429 handling is unchanged; what changed is that Aquarius no longer swallows the RateLimitExhaustedError that handling produces. Lending needed none of this and is untouched: Blend and Kinetic throw on every failed read, so no fact about our read path could reach one of their numbers in the first place.

No version bump, deliberately. Per What bumps the version, a fix that makes the implementation match the rule already documented here does not bump — the affected stored scores were wrong under this rulebook, not scored under a different one. No formula, threshold, weight or tier moves, and a dex score computed from readings is byte-identical before and after. What changes is that some cycles now publish no number where they previously published a wrong one, and risk_scores.methodology_version was never the field that distinguished those.

Every branch above is implemented and tested, and the adapter may not reinterpret it. The rules were written before the code, which is the order ../TAXONOMY.md requires; adapters/aquarius/score.test.ts now asserts each one individually, including the four that no live pool can reach — a reverting get_privileged_addrs(), a short role map, a contract-held role, and a pool whose every reserve is a wasm contract. All of them return 0, each with a detail naming what could not be assessed.

And so is the line above them. adapters/read-failure.ts decides, once, which failures are allowed to become one of those branches at all; adapters/read-failure.test.ts pins the classification and adapters/aquarius/fetch.test.ts drives the two Horizon readers against a loopback endpoint answering 404 (captured as the reading), 503, a refused connection and a persistent 429 (all three a failed run, the last still recognisable downstream as a rate limit). None of the four can be produced from live data on demand, which is why they are pinned rather than left to be discovered on the cycle they first happen.

Size floor: none, and none pending

There is no size floor in this rulebook, and there is no unwritten one waiting to be added. Gate 4 asks for a size below which a published number stops carrying information. For the two factors that ship, no such size exists — and that is a property of what they measure, not a gap left open:

  • adminKeySafety reads the same seven roles whatever the pool holds. A pool with 3 stroops in it has the same Admin, the same six single-signer keys, the same UpgradeDeadline and the same Horizon signer sets as one holding 10,000 XLM. Every input is a property of the contract and of the accounts that control it; not one of them is a quantity of anything.
  • assetControlSafety reads the same issuers. Whether Circle can freeze the pool's USDC does not depend on how much USDC is in it. A dust pool holding a clawback-enabled asset is exactly as exposed as a deep one.

Nothing here degrades as a market gets smaller, so there is nothing for a floor to protect. Lending needs one for the opposite reason, spelled out in the market-size floor: every one of its five factors falls to a can't-assess branch on an empty market and every one of those branches is 0, so an empty market would publish a danger-band number meaning the opposite of the truth. Neither dex factor has that failure mode — an empty pool's roles and issuers still read, and the number they produce is still true of it.

This is a consequence of question A's resolution, not an open item, and the distinction is the whole point. The census that would motivate a floor is real and dated — of the 148 XLM-paired pools on 2026-08-27, 18 held zero XLM and 59 held under 1 XLM, many of them 1–3 stroops, while only 16 held more than 10,000 — but a floor is needed by a size-sensitive factor, and the one that would have been size-sensitive (depthSafety) is deferred. A floor becomes necessary again the moment depthSafety lands, which is the same moment the unit of value to express one in would exist. Writing one before then would mean choosing a denomination for a factor that does not exist, in a category with no unit of value — the permanent, one-protocol decision question A declined to make.

No number is copied from lending's market-size floor or its minimum-size filter, and none may be. Those are denominated in USD against a lending pool's supplied value, which is a quantity this category cannot compute.

What this does not claim, and where an under-sized market goes. It does not claim a 3-stroop pool is worth anyone's attention — it claims that the two numbers this rulebook publishes about one are as true of it as they are of a deep pool, which is a statement about the factors and not about the market. Nothing is excluded for size, so no pool needs a coverage entry on those grounds. Which pools are registered is a separate reviewed decision about what is worth publishing, and never a statement that an unregistered pool could not be scored.


Unvalidated judgment calls

../TAXONOMY.md Gate 1 requires every threshold in a rulebook to either name the on-chain field it anchors to or be labelled an unvalidated judgment call, in both the code and methodology/, with its reasoning and its direction of error stated. This is the complete list for dex v1. Every other number in either formula names a field or is itself a reading; where one is adopted from lending rather than chosen here, that is said below rather than left to be noticed.

ConstantWhat it isDirection of error, in one line
the two weights, 0.55 / 0.45each factor's share of the overall scoreunder-weights issuer control if wrong; the gap is deliberately small because the ordering is the claim
combining the seven roles by minhow seven separately-read roles become one numberunderstates safety, never overstates it — a weak minor role binds as hard as a weak Admin
the pending-upgrade ceiling, 40what a scheduled code change does to the numbertoo harsh for a routine announced upgrade, far too lenient for a hostile one
the asset-control tier spacing, 40 and 70how far apart freeze-capable and clean-but-mutable sit40 leans lenient, 70 conservative — and min makes the lenient end bind, so this one reads generously

The two weights, 0.55 / 0.45. Argued at length in Factor weights and labelled in ../core/src/weights.ts. No external framework anchors them. If the split is wrong it is wrong by under-weighting issuer control, and a reader who holds that unmitigable-and-unannounced should dominate would push toward 0.45 / 0.55. The gap is 0.10 rather than lending's spread because the two arguments nearly cancel: adminKeySafety reaches more of the pool, assetControlSafety reaches it with no warning and no remedy.

Combining the seven roles by min. Every one of the seven can act unilaterally within its own scope, so an attacker takes whichever key is weakest and the weakest is what the number should report. The cost is that min treats the seven as equally consequential, which they are not — a SystemFeeAdmin compromise is not an Admin compromise — so a pool whose only weak key is a minor one is graded as though the code-upgrade key were weak. It therefore understates safety and never overstates it, which is the direction ground rule 4 requires a judgment call to lean. Two alternatives were considered:

  • Weighting the roles by how much each can do. It would fix the objection exactly, and it needs seven Stenion-chosen numbers where min needs none — turning a four-row type-(b) list into a ten-row one, which by TAXONOMY.md's own framing would itself be the finding.
  • Grading the owner key alone. Rejected outright. Aquarius's Admin is a 2-of-3 multisig, so this would publish a comfortable number while ignoring six live single-signer keys that can pause the market, move fees and set rewards. It is the alternative that most looks like a simplification and is in fact a decision to stop reading six of the seven readings.

The pending-upgrade ceiling, 40. It introduces no integer this document did not already publish — it is the single-master-key tier, reused on the reasoning that a scheduled code change leaves the LP holding a countdown of unreadable length rather than a signer structure. Direction of error, both ways: it is too harsh for a routine, correctly-announced upgrade, because it lowers the number of a protocol for using the very two-step mechanism this factor credits it for; and it is far too lenient if the staged code is hostile, in which case no score would be adequate. It leans conservative, and it is unexercised today — UpgradeDeadline read 0 on the router and on pools of all three types on 2026-08-29. The alternative was to leave the pending state ungraded and publish it as a disclosure; that was declined because the reaction window is the failure mode this factor's Gate 0 argument rests on being able to see, and a factor that discloses it is a role-posture factor with a footnote.

The asset-control tier spacing, 40 and 70. Only the two interior values are chosen; the ordering around them is anchored to what Stellar's flags let an issuer do.

Direction of error, and it is the one entry in this list that is not symmetric. 40 for freeze-capable leans lenient: to an LP, a freeze that is never lifted is indistinguishable from a seizure, and 40 asserts the two differ. 70 for clean-but-mutable leans conservative: it refuses to call a mutable issuer as safe as an asset with no issuer, and so understates a long-standing issuer that has never set a flag and never will. Those pull opposite ways, and the lenient end is the one that binds, because the factor is a minimum over the pool's tokens: any pool holding a freeze-capable asset is scored by the 40 and the 70 never enters the number at all. So the net lean of this constant is toward reading generously — the opposite of the role-combination rule two entries above, which understates safety and never overstates it, and the same direction as the weights' possible under-weighting of issuer control. Those two compound, and on exactly one shape of pool: one whose only weakness is a freeze-capable issuer is graded by the lenient tier and has that tier's factor weighted at the lighter of the two. That pool is where a published dex number is most likely to be too high, and it is the first thing to challenge if one looks generous.

Moving the two values together makes the factor a pass/fail on clawback alone; moving them apart makes it nearly binary on any flag at all.

Adopted from lending rather than chosen here, and labelled there: the per-account tier values (90 for an N-of-M multisig, 40 for a single master key) and the activity penalty (−3 per operation, capped at −30), from lending §3. They are the same integers applied to the same Horizon reading, and that section carries their label, their reasoning and their direction of error. They are named here so a reviewer counting this rulebook's constants finds them, and they are not re-argued here, because a second argument for the same integers is how two rulebooks start. Changing them on one side and not the other is a decision to fork them and needs its own argument — it is not an edit to one file.

Four rows, where the weight-table review was expected to bring two — and the growth is stated rather than smoothed over, because TAXONOMY.md's rule is that a long type-(b) list is itself the finding. This one is not long, and both additions are accounted for: that prediction was written when this rulebook had no formulas at all, only anchors, and the two new rows are precisely the two places a formula had to say something the chain does not state — how seven separately-read roles become one number, and what a scheduled code change does to it. Neither is a preference standing in for a measurement, and neither introduces an integer published nowhere else. On the other side the list shrank: that same prediction included rows for depthSafety's probe size, its impact-to-score mapping, and the size floor, and all three are gone because the factor and the floor are gone with them.

Three of the four are labelled by the adapter's scoring code, not by this document. The weights live in CATEGORY_FACTORS.dex and carry their label there. The other three are constants in adapters/aquarius/score.ts, which labels each one beside the branch that applies it — the per-account tiers adopted from lending by reference, the reused 40 in the open-upgrade-window branch, and the 40/70 asset-control tiers. Gate 1's "in both the code and methodology/" is therefore discharged for all four. Recorded here so the obligation stays attached to the rulebook rather than remembered.


Live ungraded state — route (c), published beside the score, never in it

Aquarius's pause surface is real, changing, on-chain and ungradable: exactly the shape operationalState exists for (Operational state). Every pool exports get_is_killed_swap(), get_is_killed_deposit(), get_is_killed_claim() and get_emergency_mode(); the router exports get_emergency_mode().

Read on 2026-08-27: router emergency mode false; two stable pools have is_killed_deposit = true — CDKVJYMN34ZIEXSLNFYHVAFF6M6FM5E2U6OHXOTBKH2WLBULXOE53YDP (XLM/AQUA, zero reserves) and CBKENQ33KITYE4JKPAWALHU4KGWV5AXQLJFUNRNIDGCRLRDENX6PYVDE (AQUA/CCKCKCPH…) — everything else false, no pool in emergency mode. The field has live variance on day one, which is what makes it worth publishing rather than a field that is always the same word.

The structural finding, and the strongest single fact about an Aquarius LP's exit risk: there is no kill_withdraw. The exported function list of all three pool wasms contains kill_swap, kill_deposit, kill_claim, kill_gauges_claim and their unkill_ counterparts — and no withdraw equivalent. Withdrawals cannot be halted by any Aquarius role. That is the same class of fact as "Blend never blocks withdrawals at any status," established the same way: it is a property of the deployed code, not a promise.

Read the state through the getters, never through instance storage. The three pool types disagree about key names — standard writes IsKilledClaim, concentrated writes ClaimKilled/IsKilledSwap/EmergencyMode — and a flag that has never been toggled has no key at all. The getters normalise all of it; raw storage would make an untouched pool indistinguishable from one that was read wrongly.

The swapDisabled rung, and why the shared ladder gained one

dex registers its own operation vocabulary — { swap, deposit, withdraw, claim } — and its own canonical ordering and classifier in ../core/src/operational-state.ts, beside lending's. No lending operation was renamed, added or removed; renaming one would rewrite blocked in every stored operational_state and every API response.

OperationalLevel stayed one shared ladder, and gained a rung rather than forking — an open question when the category was admitted, recorded here so it is not re-litigated:

A pool with is_killed_swap = true but deposits and withdrawals live would have classified as active under the ladder as it stood — true about exit, and wrong about the market. blocked: ['swap'] carried the fact, but level is the field a reader scans, and a dead market reading "active" is the quiet misstatement the whole type exists to prevent. So the ladder gained swapDisabled: cannot trade; depositing and withdrawing both work.

Why not generalize borrowingDisabled into one shared coreActivityDisabled instead, which would have been the more honest name for both: borrowingDisabled is a value stored in every historical operational_state and published on the public API, so renaming it is a breaking change to stored data and to the API contract, inside a change whose entire purpose was to add a category. Two named rungs and an accurate level beat one elegant rung and a migration nobody asked for. The two sit at equal severity and can never meet, since blocked carries one category's vocabulary.

exitDisabled is unreachable from Aquarius — there is no kill_withdraw — and the rung stays, because it is the top of the ladder every category is measured against and a second DEX that can freeze withdrawals must have somewhere to say so. claim gets no rung at all: an LP whose reward claim is killed can still withdraw every unit of principal. It stays in blocked because it is true, on the same reasoning lending gives for repay and liquidate.

What was rejected, and why it must not be re-proposed

Gate 3. Every candidate that was considered and dropped, by name, with its reason — and with the measurement where one was taken.

liquidity_calculator.get_liquidity() as a size or depth measure

Aquarius ships its own cross-pool liquidity metric (CCUKQWLM…, get_liquidity(pools) -> Vec), used to allocate AQUA rewards. It is the obvious thing to reach for, it is on-chain, and it clears Gate 8 — and it is not comparable in any economic sense. The evidence is a single pool: on 2026-08-27, CBMSBM6EABGBNZ47WZTLX6WJOB3ETO2GNPQYC2FX6LJXUL7TRQFZ3IA3 — a standard XLM/DATAVAULT pool holding 32,733 stroops, i.e. 0.0032733 XLM, roughly a third of a US cent — ranked 4th of all 340 pools by this metric. It is dominated by raw token quantities, so a pool paired against a huge-supply token outranks the XLM/USDC book. Fails Gate 1 (no economic anchor), and would make any Gate 4 floor built on it meaningless.

LP concentration

Named in ../ROADMAP.md as a DEX candidate. Fails Gate 8 outright. The LP share token is a plain SEP-41 contract (CAVKLYY4… for the XLM/USDC standard pool, wasm 07bab30e…); its complete exported interface is balance / transfer / transfer_from / approve / allowance / mint / burn / burn_from / decimals / name / symbol / upgrade — there is no holder enumeration and no holder count. Recovering the distribution needs an indexer over transfer events, which is the off-chain data path Gate 8 disqualifies however well-anchored everything else is. Concentrated-liquidity positions are worse: keyed by (owner, tick_lower, tick_upper) through get_position(...), with no way to enumerate owners at all.

Price divergence from a reference market

Also named in ../ROADMAP.md. Rejected on the reasoning ROADMAP.md already records for market-depth-aware oracle scoring: an AMM's price legitimately differs from SDEX by up to the fee plus whatever arbitrage has not run yet, so a divergence reading cannot separate a real problem from a normal one. Horizon order books are also trivially spoofed with walls that are never hit. A cross-pool reference is worse still — it is another market standing in for the one being scored, which is the substitution Gate 8 exists to refuse.

liquidity_pool_plane as a source for any scored input

The plane (CCABO2IQ…, get(pools) -> Vec) returns type, init args and reserves for many pools in one call, and is tempting as the bulk read. It is a cache written by the pools (update(pool, pool_type, init_args, reserves), with ReservesSyncLedger recorded on the pool), so it can lag the pool it describes. Acceptable for discovery; the pool contract is the source for anything that reaches a number.

Fee tier as a factor

get_fee_fraction() varies genuinely — standard pools use only 10/30/100 bps, but stable pools were found at 1, 5, 10, 15, 22, 25, 30 and 50 bps. A higher fee is worse execution, not a failure mode; grading it would dress a pricing preference as a risk measurement. It belongs inside depthSafety's round-trip cost, where it already is, and nowhere else.

Reserve imbalance against the pool's target ratio

Real for a stable pool, where drift from 1:1 is a genuine de-peg signal — and meaningless for standard, where imbalance is the price. A factor only 42 of 340 markets can be graded on is a per-market rulebook, which ground rule 1 forbids.

Aquarius's own AMM API (amm-api.aqua.network)

Published in their documentation, and it would answer several of the questions above directly. Gate 8. No further discussion.

Concentrated-liquidity-specific factors

Tick distribution and active-liquidity fraction are readable — get_active_liquidity(), get_slot0(), get_tick(), get_chunk_bitmap_batch() — and genuinely interesting. 26 of 340 pools have them. A factor only those pools can be graded on is a per-market rulebook, same as reserve imbalance. Revisit only if concentrated pools become the majority of pools.

Question A — resolved: no depth factor until there is a unit of value

Decision: option 4. dex ships with two factors, adminKeySafety and assetControlSafety. There is no depthSafety, no size floor and no denomination logic, and none may be added without a further published decision. The problem this resolves, and the three options declined, are recorded below so none of it is re-proposed from scratch.

It was decided on reversibility, which is the argument that outranked the others. Adding a factor to a live category later is additive: it lands as a labelled version bump, old scores stay readable as what they were, and the discontinuity is published rather than hidden. Walking back a published depth denomination is not symmetric — it would mean revising what stored scores meant, and this project's own rule is that history is never backfilled across a bump and cannot be, since risk_scores keeps only outputs and no row can be recomputed. So the two directions carry very different costs: shipping without depth is a gap that can be closed, while shipping the wrong depth denomination is a mistake that cannot be undone. Under that asymmetry the thin category wins.

Gate 4 is satisfied by not needing a floor, rather than by having one — and this is a consequence of the decision, not a dodge. Lending's floor exists because an empty market publishes 0 in the danger band, meaning the opposite of the truth. Neither surviving factor is size-sensitive: a pool holding 3 stroops has exactly the same seven admin roles and exactly the same token issuers as one holding 10,000 XLM, and both readings are equally true of it. There is no quantity here that stops carrying information as a pool gets smaller, so there is nothing for a floor to protect. A floor becomes necessary again the moment depthSafety lands, which is the same moment the unit of value to express it in would exist.

The problem itself, unchanged, because it is what any future attempt has to solve:

Aquarius reads no oracle, so nothing in its own contracts values a pool. Two consequences:

  • This category's size floor has no denomination. A floor is badly needed: of the 148 XLM-paired pools, 18 hold zero XLM and 59 hold under 1 XLM (many hold 1–3 stroops); only 16 hold more than 10,000 XLM. Over half the registry would otherwise publish a number computed from dust. But "below what?" has no answer in Aquarius's own terms.
  • depthSafety is degenerate for 272 of 340 pools if the trade is sized relative to the pool. For a constant-product pool, the price impact of swapping a fixed fraction of the input reserve is a closed form that does not contain the reserves — so every standard pool at the same fee tier scores identically and the factor collapses into the fee tier it was told not to grade. It would discriminate only among stable (amplification) and concentrated (tick distribution) pools. An absolute trade size fixes this, and needs the unit that does not exist.

No number in this document is copied from lending's market-size floor or its minimum-size filter, and none may be: those are denominated in USD against a lending pool's supplied value, which is a quantity this category cannot compute.

The options, with gate verdicts. None is adopted here:

#OptionVerdict
1Denominate in XLM, pool-relative. 148 of 340 pools contain XLM directly, so no price is needed for those.Clears Gate 8. Leaves 192 pools unscorable and needs a coverage status for them.
2Derive prices from Aquarius's own reserve ratios along an XLM-paired path.Clears Gate 8 on the letter. Almost certainly fails Gates 1 and 2: a price computed by Stenion from a manipulable pool is a fabricated anchor — precisely what lending's rejected "Stenion-computed deviation" candidate already refused.
3Anchor to Aquarius's own declared cost of existence — get_standard_pool_payment_amount() = 300,000 AQUA, the price the protocol itself charges to create a pool, and the closest structural analogue to Blend's PoolConfig.min_collateral.Clears Gate 1 as a type-(a) anchor. Still needs option 1 or 2 to compare a pool's reserves against it.
4Score no depth factor at all; publish depth as a route-(a) disclosure and admit dex on adminKeySafety + assetControlSafety.Honest, but a two-factor category is thin enough that Gate 0 should be re-argued before accepting it.

Why option 2 was rejected, even though it is fully on-chain

Deriving prices from Aquarius's own reserve ratios along an XLM-paired path clears Gate 8 on the letter — every read is Soroban RPC against the protocol's own contracts, with no aggregator and no off-chain source anywhere. It is rejected anyway, and the reason is worth stating precisely because "it's all on-chain" is exactly what makes it tempting.

A reserve ratio is not a price. It is spot state, and it is cheap to move. A large one-sided swap — or a flash loan, which needs no capital at all — skews a pool's ratio for exactly as long as it takes us to read it. So the "price" would be whatever an adversary wanted it to be at the instant of the read, and the factor built on it would report a comfortable number precisely when someone was active in the pool. A score derived from the same system it is meant to be checking goes blind at the only moment it matters.

This is the same failure shape as the price-deviation candidate already rejected for lending's oracleSafety (see index.md §2's rejected candidates) — a Stenion-computed figure standing in for an anchor the protocol does not publish — and it fails two gates rather than one:

  • Gate 1: the resulting number is not a protocol-declared anchor. It is Stenion-computed from mutable state, which is the definition of an unanchored threshold wearing an anchor's clothes.
  • Gate 2: it does not fail to the unsafe end. It fails by lying confidently — publishing a plausible, precise, wrong number, which is worse than publishing nothing and much worse than publishing 0.

A TWAP-based approach is deferred, not dismissed. Time-averaging would genuinely resist the single-block skew above, and the objection in this section does not defeat it. It is a materially different and harder design — it needs a window, an update cadence, a manipulation-cost model for that window, and a story for pools that trade rarely — and none of those has an anchor in Aquarius's contracts either. Designing it against exactly one protocol would bake Aquarius's shape into a category rulebook. Revisit when a second dex protocol exists to calibrate against.

Why options 1 + 3 were not taken either, despite being the submission's own lean

The original recommendation was XLM-relative sizing plus an absolute floor anchored to the 300,000 AQUA pool-creation cost. It is the strongest of the three declined options and it was still declined, on three grounds:

  • It is not reversible the way option 4 is. A published depth denomination becomes what stored scores mean. Changing it later requires revising history, which the no-backfill rule forbids — so the choice would be effectively permanent, made now, on one protocol's data.
  • It strands 192 of 340 pools on the category's flagship factor. Only 148 pools contain XLM directly, so the rest would sit in coverage-only limbo — unscored on the very measurement the category was built to publish. A flagship factor that does not apply to 56% of the market is a weak flagship.
  • The floor does not answer the question it appears to answer. "It cost 300,000 AQUA to create this pool" is a fact about the protocol's fee schedule, not about whether the pool is deep enough for a number computed from it to mean anything. It clears Gate 1's anchoring requirement — the figure really is protocol-declared — while leaving the actual depth-adequacy judgment unstated and ungraded underneath it. An anchor that anchors the wrong quantity is a worse failure than an admitted judgment call, because it looks anchored.

Both are worth revisiting once a second dex protocol exists to calibrate against. The objection throughout is not that these ideas are bad; it is that every one of them has to be designed against a single protocol today, and a category rulebook designed against one protocol is that protocol's rulebook wearing a category's name.

What holds until then: no adapter may implement depthSafety; no size floor exists, and none is needed by the two factors that ship; depth is published as a route-(a) disclosure rather than graded.

Comparability: within dex yes, across categories no

Gate 5, stated in both directions, because only one of them is obvious.

Within dex, two safetyScores are comparable. Every protocol scored under this rulebook is graded by the factors above, with the same formulas and the same thresholds, under the same version stamp — ground rule 1, which binds every adapter in a category with no exceptions. That is what makes the ranked block a ranking rather than a list.

Across categories they are not, and nothing may present them as if they were. A dex safetyScore of 70 and a lending safetyScore of 70 were produced by different factor sets, different formulas and different weights, from data that has no quantity in common. Neither number is evidence about the other; "the DEX is safer than the lending market" is not a statement either score supports.

This is enforced, not merely stated:

  • buildRegistryView (dashboard/app/lib/registry-query.ts) publishes RankedCategoryGroup[] and no flat ranked array, so each category is its own block numbered 01..n within itself and there is nowhere for a cross-category ranking to live. Name sort is the sole ordering allowed to merge categories, because alphabetical asserts no ranking.
  • Every entry carries its category through to the board and to the API responses.
  • risk_scores.category is stamped beside risk_scores.methodology_version on every run, because both counters start at 1 and the integer alone does not identify a rulebook.

Sharing the adminKeySafety key between the two categories changes none of this. The key names the question; it does not claim the answers were computed the same way — and the formulas above are the proof: lending grades one admin account's signer set, this grades seven roles' worst posture under a pending-upgrade ceiling. See factor 1.

One adapter, many markets

The same rule lending already runs under, and Aquarius has the same shape as Blend: exactly one wasm per pool type, each matching the hash the router itself declares — ConstantPoolHash ae0da5a8…de9852 (272/272 pools), StableSwapPoolHash f1077e0b…e747cd (42/42), ConcentratedPoolHash 12fca5a7…d37ee6 (26/26). Verified against the router's own declared hashes rather than assumed.

So a second Aquarius market is a config entry and no new scoring code. CLAUDE.md's "one adapter may serve several markets; a market never gets its own adapter" applies here unchanged, and nothing on a per-pool config may be a threshold, weight or formula — that would be a per-pool rulebook.

Operational state is published, never scored

The decision, up front: pause/frozen state is a published field beside the score, and it is deliberately not a factor, not a multiplier, and not any input to a number. It was decided this way on 2026-08-25 after reading both protocols' contracts, and this section exists so it is not re-litigated by the next person who notices a paused pool with an unchanged score.

Both adapters had always read a pause signal — Blend's PoolConfig.status, K2's router.is_paused() — and neither had ever used it. That could not stay true indefinitely: every adapter written while it stayed unresolved would have to be retrofitted later.

What each protocol actually means by "paused"

Read from the contracts, not from the documentation — the published docs name Blend's states but publish no numeric mapping, and the mapping circulating in search results is partial and partly wrong.

Blend V2 gates in require_action_allowed (pool/src/pool/pool.rs), which is the entire rule:

if (status > 1 && (action == 4 || action == 9))     // Borrow, DeleteLiquidationAuction
|| (status > 3 && (action == 2 || action == 0))     // SupplyCollateral, Supply
{ panic!(InvalidPoolStatus) }

RequestType numbering is from pool/src/pool/actions.rs; the setter paths are execute_set_pool_status (admin) and execute_update_pool_status (permissionless) in pool/src/pool/status.rs.

statusBlend's nameBorrowSupplyWithdraw / Repay / liquidation fillsWho can set it
0Admin Activeyesyesyesadmin only (needs the backstop threshold met and Q4W < 50%)
1Activeyesyesyespermissionless only (backstop healthy)
2Admin On-Icenoyesyesadmin only
3On-Icenoyesyeseither — admin, or automatically at Q4W ≥ 30% / below the backstop threshold
4Admin Frozennonoyesadmin only; supersedes the backstop, which cannot move it
5Frozennonoyespermissionless only — automatically at Q4W ≥ 60% (≥ 75% from status 2)
6Setupnonoyesinitialization only; supersedes everything

Blend never blocks a withdrawal or a repayment at any status. The only user-facing action blocked below the supply threshold is cancelling an in-flight liquidation auction, which is a wind-down-safely posture rather than a restriction on depositors.

K2 (Kinetic) has two layers. storage::is_paused is checked at the top of validate_supply, validate_withdraw, validate_borrow, validate_repay and validate_liquidation, and again in the flash-loan and two-step liquidation entry points — so a paused K2 halts everything, withdrawals included, and deposited capital cannot leave. Separately, each reserve carries its own gating flags in the ReserveConfiguration bitmap it already publishes its decimals in (contracts/shared/src/utils.rs, bits 50–53): active, frozen, borrowing_enabled, paused. A cleared active or a set paused blocks every operation on that reserve; frozen blocks supplying and borrowing while leaving withdrawals open; a cleared borrowing_enabled blocks only borrowing.

The shared representation

Because those two vocabularies do not map onto each other, the published state is named by what is blocked, which is the one axis on which the protocols are genuinely comparable:

LevelMeaningBlendK2
activenothing restrictedstatus 0, 1not paused, every reserve open
borrowingDisabledcannot borrow; supply and exit both workstatus 2, 3reserve borrowing_enabled = false
entryDisabledcannot borrow or supply; existing positions can still exitstatus 4, 5reserve frozen
exitDisabledcannot withdraw — capital cannot leaveunreachablerouter.is_paused(), or reserve paused/!active
notOperationalthe market was never openedstatus 6no analogue

Where a protocol gates per reserve, the most restricted reading is published — the same worst-reserve convention §2, §4 and §5 use, for the same reason. Alongside the level: the protocol's own reading verbatim (PoolConfig.status = 4), the exact operations blocked, when it was read, and whether the value is one only an admin could have set. That last field is indeterminate for Blend's status 3, because execute_set_pool_status accepts it too — reading even/odd as "who did this" would be right six times in seven and wrong on the one value where it matters.

notOperational is reachable in principle and not in practice: every Setup pool in the 2026-08-22 factory survey held exactly $0.00 and is already excluded by the market-size floor.

Why it is not scored

Three options were weighed — a sixth factor, a multiplier on the overall score, and a published flag. The flag was chosen, on four grounds:

  1. No on-chain datum resolves the ambiguity. A pause can be an admin containing a threat or an admin abandoning a market, and neither protocol's state carries a reason. Distinguishing them needs off-chain announcements, which adapters may not read. A "context-dependent" factor with nothing to condition on is a flat penalty in costume, and any magnitude for it would be invented — this document's standard is that a threshold is anchored to a protocol's own parameter or labelled an unvalidated judgment call, and there is no anchor here at all.
  2. The one axis clean enough to score does not exist on both protocols. The strongest scored variant was not a sixth factor but folding exitDisabled into liquiditySafety: that factor is defined as the withdrawal cushion, and a cushion you are contractually barred from drawing on is zero by definition rather than by judgment — no invented magnitude required. It was rejected anyway, because exitDisabled is structurally unreachable on Blend. A rule that is live code on one adapter and dead code on the other satisfies ground rule 1 in form only. Recorded here so it is not re-proposed as the obvious fix.
  3. A scored rule would have shipped untested. Every registered market was fully operational when this landed — Blend status 1, YieldBlox status 0, K2 unpaused with all four reserves open — so the before/after comparison a scored change requires would have compared each number against itself. What a factor would have done is worse than nothing: "not paused" scores 100, so a sixth factor at weight w raises every active protocol's score by (100 − score) × w, handing Blend, K2 and YieldBlox 5–11 free points for the ordinary state. That is the five-way redistribution the weights note already declined once.
  4. A multiplier would break the score model. safetyScore is one published line, and this document's worked example spells the arithmetic out. A client that fetches factors and reproduces the score would stop getting the same number — a verifiability platform whose published factors no longer reconstruct its published score has traded away more than the change buys.

There is also no decay problem, which both scored options have and neither answers: snapping back on unpause puts a step in the history chart that is not a change in risk, and a cooldown invents a time constant from nothing. A published state is a live reading — correct at every instant, with nothing to recover from.

What this costs, stated plainly. A reader who looks only at the number is not protected by the flag. That is why the flag is a first-class field on both the leaderboard and the detail response rather than a footnote, and is rendered beside the name and score everywhere either appears — the same treatment, for the same reason, as deployedOn. If it ever stops being rendered there, the decision not to score has quietly become a decision to hide.

No version bump. Nothing here changes a formula, a threshold or a weight, and no stored score moves, so lending's methodology version stays at 1. That the state cannot reach a factor is enforced rather than intended: adapters/blend/score.test.ts and adapters/kinetic/score.test.ts each assert a byte-identical factor map across every restricted state their protocol can be in. If pause state ever moves a number, those tests fail before the change ships.


Findings are published, not scored — and how they must be written

Verifiable observations we can't or won't grade go in the protocol page's Findings section (dashboard/app/lib/protocol-notes.ts), never into a factor. Nothing there is read by any scoring path, and a note — favourable or not — can never move a number.

Findings are the STATIC half of ungraded publication. They are hand-written and reviewed in a PR, which is right for an observation someone had to go and establish, and wrong for a reading that changes every five minutes. The live half is operational state: measured every cycle, published as a typed field, and equally never graded. A new ungraded observation belongs in whichever of the two matches how it is obtained — never in a factor, and never invented as a third mechanism.

A note must survive the history it was drawn from. Twice now a Findings note has outlived the stored runs behind it: once when the development-era history was discarded, and again when the briefly-live v2 rows were. Score history is not an archive — it is discarded across a rulebook change and cannot be recomputed, because risk_scores keeps only outputs. A note written as "our history shows X" therefore decays into an unverifiable claim on a page whose entire pitch is that you don't have to trust us.

So every note citing our own observations follows the same form:

  1. Cite a closed window, with both ends stated. "Between 2026-08-11 18:16 and 2026-08-18 15:55 UTC, 1,469 runs" — not "93% of runs", which silently means something different every time the cron fires. A reader re-running the query later must be able to tell that a different number is a later window, not a contradiction.
  2. Say the counts are a snapshot of that window and do not update.
  3. Phrase the underlying claim so it stays checkable from chain after the history is gone. Our runs are evidence that a condition persisted; the condition itself must be one anyone can observe today, directly from the contracts. If the only support for a claim is rows in our database, it is not a finding — it is an assertion.
  4. Give the exact verification steps — contract, method, field, and what to compare against. If we can't say how a reader would check it themselves, it doesn't go in.
  5. Claim only what was measured. Where a sub-signal wasn't recorded separately, say so and scope the claim to the runs that carry it, rather than generalising across all of them.

Disputing or changing a threshold

Every number in this document is meant to be challengeable — especially the ones labeled "unvalidated judgment call." If you believe a threshold, weight, or formula is wrong (including if you are a protocol being scored):

  1. Open a GitHub issue against this repository describing the specific threshold/formula and why you think it's wrong. Anchor your argument to something external where possible (a protocol's own on-chain parameter, a published risk framework, observed data) rather than preference.
  2. Or open a pull request editing this file directly with the proposed change and its justification. A change to methodology/ must be accompanied by the matching change to the adapter code (and vice versa) — the two are not allowed to drift.
  3. Maintainer review is required, at the same bar as adapter code changes. A methodology change affects every protocol's number, so it is reviewed at least as carefully as a code change — not merged on preference, and never merged because a scored party requested it. Per the ground rules above, no change is ever accepted in exchange for payment.

Changes that alter what a factor means (e.g. adding or removing a factor) are breaking changes to the shared taxonomy in core/src/types.ts and are held to a higher bar again — they affect every adapter at once.