Methodology
Every formula, threshold, and weight Stenion uses — extracted directly from the shipped adapter code. It exists so anyone, including the protocols being scored, can verify and challenge the rules, not just the output. Code and this document are not allowed to drift.
View source on GitHubStenion Scoring Methodology
This document is the source of truth for how every safety factor is calculated.
It exists so that anyone — including the protocols being scored — can see, verify, and
challenge the actual rules, not just the output numbers. Every formula below is extracted
directly from the shipped code (currently adapters/blend/ and
adapters/kinetic/); this file is not a summary of intent, it is
the rulebook the adapters must implement.
One formula per category, per-protocol data sources. Every factor's formula, scale, and
thresholds are fixed here and identical across every protocol in a category — and across
markets: the three Blend pools Stenion scores run one adapter and one rulebook, differing only
in the pool each reads. Each category owns a section below, holding its own factor list, weight
table, worked example and version changelog. There are two — lending, which every scored market
runs under, and dex, published but not yet scoring anything. What legitimately
differs per adapter is only where the raw inputs are read on-chain — e.g. Blend reads a per-reserve max_util cap,
while Kinetic (K2), being Aave-V3-style, has no such cap and instead anchors the same
utilization formula to its own OPTIMAL_UTILIZATION_RATE (see §5). The anchoring pattern
("grade against the protocol's own on-chain parameter") is the invariant; the specific
parameter that pattern resolves to is a documented per-protocol fact, not a new threshold.
If the code and this document ever disagree, that is a bug — open an issue (see Disputing or changing a threshold).
Current version
Versions are per category, on independent counters that each start at 1, so a version number alone does not identify a rulebook — the category and the number together do. One row per category, and the changelog behind each lives in that category's own section:
The two 1s are not the same 1. Counters are independent and each starts at 1, so dex v1 and
lending v1 are two different rulebooks rather than two editions of one, and neither is older than
the other. The category and the number together identify a rulebook; the number alone does not.
Dex, methodology v1 — complete, and in use since 2026-08-29. The two-factor AMM rulebook in
Dex — adminKeySafety weighted 0.55 and assetControlSafety weighted 0.45, each with
its formula, its thresholds and its worked example. One market is scored under it —
aquarius-xlm-usdc, read by adapters/aquarius/ — and every stored dex row carries version 1.
The factor set was admitted first and the weight table reviewed separately, because a weight can
only ever be an unvalidated judgment call and arguing one alongside the factor set would have buried
it; both are version 1, since no score was ever published under the factor set alone. A third factor, depthSafety, was
proposed and deferred: Aquarius publishes no unit of value to denominate a trade size in, and
every way of inventing one is either a fabricated anchor or a permanent choice made on one
protocol's data. No size floor exists and none is pending — neither surviving factor is
size-sensitive. It was published ahead of the adapter, which is the order
../TAXONOMY.md requires: a category's rulebook must be reviewable before an
adapter is written against it.
Lending, methodology v1 — the rulebook described in the Lending section, in full,
including oracleSafety scoring both price freshness and manipulation resistance (§2), the
minimum-size filter §4 and §5 select reserves through
(The minimum-size filter), and the
market-size floor that decides whether a market is scorable at all.
The floor is a precondition rather than a formula — it moves no number and did not bump this
version.
The rest of this section — what bumps a version, and what a boundary means for a stored score — is policy that applies to every category, not just lending.
Versioning begins here. methodology_version = 1 is the only version lending's rulebook
defines, and the only one any stored row will carry. A version 2 was briefly live in the code
— stamped onto runs between 2026-08-14 11:25 and 2026-08-18 11:30 UTC, before the rulebook was
flattened back to v1 — and that history is discarded rather than migrated, for the same reason
the development-era history was: it was computed under a rulebook that no longer exists, and
nobody was downstream of it. After that discard there is no v1-versus-v2 boundary in
risk_scores, and none to look for.
Earlier development history was discarded, not migrated
Before this point Stenion accumulated a few weeks of scored runs during development, under earlier iterations of these rules. That history was deleted rather than carried forward, and this is a deliberate, recorded choice rather than a silent one. Three reasons, stated plainly:
- It contained scores computed under two known bugs, since fixed. Those numbers were wrong under their own rulebook, not merely scored under a different one.
- It predates the oracle robustness work (§2), so its
oracleSafetyvalues measured price age alone — a signal we now consider misleading rather than merely incomplete. - Nobody was downstream of it. Every row came from our own cron during development; no external consumer had been built against the API, and the only reader of the history was our own score chart. Marking a discontinuity in a dataset nobody had read would have been bookkeeping, not disclosure.
A clean history starting from a rulebook we actually stand behind is more honest than a marked-up one carrying forward numbers we know were wrong. This is the last time that reasoning applies. From here on, history is never deleted and never backfilled — the version stamp exists so a change is labeled instead.
What bumps the version, going forward
Bump when a change alters what a number means — a factor starting or stopping measuring something, a threshold's anchor changing, a re-weighting, or any formula change that moves scores for unchanged on-chain state. Don't bump for a fix that makes the implementation match the rule already documented here (the stored scores were wrong, not scored under a different rulebook — say so in the changelog instead), for adding a protocol or an adapter, or for wording, disclosure, and presentation changes. The test is simple: if comparing an old score to a new one would mislead, bump; if the old score was just incorrect under this same rulebook, don't.
Scores across a boundary are not comparable
No boundary survives in the stored data — the v2 rows above were discarded — but the machinery that marks one is live and tested, because the first real bump must be legible on the day it happens rather than built in a hurry then:
- The indexer stamps
risk_scores.methodology_versionfrom this category's entry inMETHODOLOGY_VERSIONS(core/src/category.ts) at write time, resolved from the target's own category. An adapter has no say in the version; it only declares which category it belongs to.risk_scores.categoryis stamped beside it, because every category's counter starts at 1 and the pair — not the integer — identifies a rulebook. - The score-history chart on each protocol page breaks the line at a version change rather than drawing through it, and the run list labels the break. Both paths are covered by fixture tests, since live data cannot exercise them.
- The version is returned on every history point and on the protocol detail from
GET /api/v1/protocol/:id. To check which rulebook produced a stored score, read that column — don't infer it from the date.
History is not backfilled across a bump, and cannot be — risk_scores stores only outputs
(the score and the factor map), never the raw on-chain inputs a run was computed from, so no
one, including us, can recompute an old row under new rules.
Ground rules (non-negotiable)
- The same formula applies to every protocol in a category, with no exceptions. A
factor's formula, its weight, and its thresholds are fixed in that category's section
below, in one shared place. They do not vary per protocol, per adapter, or per anything
else within the category. A category boundary is the one and only place a rule may
differ — and it differs because the factor sets differ, not because a protocol asked:
utilizationSafetyweighted 0.20 says nothing about an AMM with no borrow cap. That is a different rulebook, published in full under its own heading and versioned on its own counter, never lending's rules bent to fit. Two protocols in the same category are always graded by the same rules. - Payment never changes a threshold or a formula. Protocols can pay for visibility, speed, or private tooling — never for a better number. A paid tier cannot move a threshold, reweight a factor, or alter a curve. The only thing that changes a protocol's output is its own real, on-chain data.
- Different protocols can and should score differently. That is the point. What must never differ is the rule being applied. Blend scoring 54 and a hypothetical protocol scoring 80 is a result of their data, not of two different rulebooks.
- No fabricated numbers. Where real data genuinely isn't available for a factor, the score uses a clearly-flagged neutral baseline (called out explicitly below) — never an invented, plausible-looking value.
- AI never sets a score. Any AI feature only explains or summarizes the numbers these formulas produce. It never generates an independent risk assessment.
Score model
-
Overall score: 0–100, higher = safer. API/field name
safetyScore. -
Every factor is on the same scale: 0–100, higher = safer. Factor names end in
*Safetyso a name never disagrees with its number — acollateralSafetyof 70 means well-diversified (safe), not "70% concentrated." -
The overall score is a weighted mean of the category's factors, renormalized over whichever factors are non-null (so a genuinely inapplicable factor doesn't drag the score toward zero rather than being excluded):
safetyScore = round( Σ(factor.value × factor.weight) / Σ(factor.weight) )
This arithmetic is shared; the factors it averages are not. The formula above is
category-agnostic — it reads a value and a weight and nothing else — and it is implemented once,
in scoreFactors (core/src/scoring.ts), which every adapter of every
category calls. Which factors exist and what each is weighted is per category: declared
once in CATEGORY_FACTORS (core/src/weights.ts) and published in that
category's own section below. So the weight table and the factor list live under
Lending, not here — a second category would bring its own, not edit lending's.
A factor may publish a components breakdown — the sub-signals behind its value.
Components with a numeric value are what the factor was computed from; components with
a null value are disclosures: real, readable on-chain quantities we publish but
deliberately do not grade, because scoring them would invent comparability the data does
not support (see §2c and §2d). A null component is never missing data.
One published field sits outside this formula entirely: operationalState, which reports
which user operations a market's own contracts are currently refusing. It is not a factor, not a
multiplier, and not an input to safetyScore — see
Operational state is published, never scored for
why, and for why that is a decision rather than an omission.
Lending
Everything from here to the Dex heading is lending's rulebook and lending's alone
— its version changelog, its factor weights, its worked example, and the five factors themselves.
It is one of two categories Stenion publishes a rulebook for (PROTOCOL_CATEGORIES in
core/src/category.ts), and it is the only one anything is scored under
today. This section was written as the shape a second one would take — its own heading, its own
changelog, its own weight table, its own factor list — and Dex is that second one,
taking exactly that shape. Nothing above this line is lending-specific; nothing below it may be
assumed to hold for a category that isn't lending, and in particular a dex score is not
comparable with a lending one.
Which protocols are scored under it: every market on the registry today — the three Blend
pools (adapters/blend/) and Kinetic/K2
(adapters/kinetic/). Ground rule 1 binds all of them to what follows.
Version changelog
Every scored run is stamped with the rulebook version that produced it
(risk_scores.methodology_version, from the lending entry in METHODOLOGY_VERSIONS in
core/src/category.ts), and it is surfaced on the API's protocol detail
and on each history point. Versions are per category and counters are independent, so this
changelog is lending's; another category's v1 is a different rulebook, not an earlier one.
Each category gets exactly one such table, in its own section. The changelog:
| Version | Effective | Change |
|---|---|---|
| 1 | initial | The five-factor model as documented here. oracleSafety scores price freshness and manipulation resistance (§2); liquiditySafety/utilizationSafety score only reserves clearing the minimum-size filter (§4). |
One row, and that is the point: v1 is where versioning starts, not where it started counting again. Development-era history under earlier iterations of these rules was discarded rather than migrated — what that was and why it was deleted is at the top of this document, along with what does and doesn't warrant a bump. See Current version.
Corrections that did not bump the version
Fixes where the implementation disagreed with this document and the document was right. The rulebook did not change, so these are not version boundaries and stored scores remain comparable across them — but they are recorded here rather than left silent, because a score did change shape even if no published number moved.
The entry below was verified against the development-era history that has since been discarded (see Current version), so its row counts are no longer re-checkable. It is kept as the record of a correction, not as a live claim about stored data. The fix itself is in the shipped v1 rulebook.
| Date (UTC) | Correction |
|---|---|
2026-08-16 | liquiditySafety (§4) and utilizationSafety (§5) returned 100 when no reserve qualified for their minimum, in both adapters. Both are a minimum over a filtered set of reserves; over an empty set that is undefined, not the top of the scale — so an unassessable pool published "maximally safe" from no data, contrary to ground rule 4. Both now return 0, matching collateralSafety's existing treatment of the same case. No published score was affected: the path had never executed — verified by scanning the entire stored history of both protocols for the signature the defect leaves in a factor's detail (a worst reserve (…) naming no asset, since worstAsset stayed empty when nothing was measured), with zero matches. Re-checked on 2026-08-16 against 1,923 rows; liquiditySafety has ranged 20–34 and utilizationSafety 10–18 across that history, never approaching the 100 the empty path would have published. |
Amendments folded into v1
Changes to the rulebook made while v1 was still being finalized as the comparability baseline — before any surviving stored history existed. These are not version boundaries: v1 is defined as the rulebook this document describes, and these are part of that definition rather than a departure from it. None of them left a step in a published history, because the history they predate was discarded rather than carried forward. Each gets a row anyway, so that what v1 means is traceable rather than assumed: "no bump required" does not mean "no record required".
This section closes when the next change lands. From that point the rule in What bumps the version applies without exception, and a change that alters what a number means bumps to v2.
| Date (UTC) | Amendment |
|---|---|
2026-08-18 | §4 and §5 gained the minimum-size filter: both now select the worst reserve only among reserves clearing the protocol's own declared minimum exposure or 0.5% of the pool's supplied USD. This changes what the two factors measure, which is why it is recorded rather than treated as a correction. No live score moved when it landed — verified against both protocols on the day: Blend excludes nothing (its smallest reserve is ~$3.4M against a $5.00 min_collateral), and K2's dust reserve had already stopped being its worst reserve. What it changes, measured on the frozen 2026-08-16 snapshot: liquiditySafety 34 → 44 and utilizationSafety 18 → 30 (score 24 → 28), by excluding a $3.00 reserve holding 0.19% of a $1,571 pool. |
History is not backfilled across a version bump, and cannot be. risk_scores stores
only outputs — the score and the factor map — never the raw on-chain inputs a run was
computed from, so an old row cannot be recomputed under new rules by us or by anyone. The
discontinuity is real and permanent; the version stamp exists so it is legible rather than
appearing as an unexplained step in a chart.
Factor weights
Lending's weights, and lending's only. This table is the published face of
CATEGORY_FACTORS.lending in core/src/weights.ts, which is where the
adapters read them from — neither adapter contains a weight of its own, and
core/src/scoring.test.ts parses this table and fails if the two disagree in either direction.
| Factor | Weight |
|---|---|
oracleSafety | 0.25 |
collateralSafety | 0.20 |
adminKeySafety | 0.20 |
utilizationSafety | 0.20 |
liquiditySafety | 0.15 |
| Total | 1.00 |
Worked example (live Blend Fixed V2 pool, 2026-08-14, lending methodology v1):
70×0.20 + 100×0.25 + 40×0.20 + 22×0.15 + 14×0.20 = 53.1 → 53.
Weights are an unvalidated judgment call, not an external fact.
oracleSafetycarries the most weight because an untrustworthy price silently poisons every other measurement — collateral value, utilization and liquidity are all priced off it. Liquidity carries the least because it partly overlaps utilization. There is no external framework these exact weights are anchored to yet — they are open to challenge like any threshold below.Oracle robustness was folded into
oracleSafetyrather than given its own factor, partly for this reason: a sixth member would have forced a redistribution across all five, layering a second unanchored judgment call on top of one already flagged as unanchored. The taxonomy incore/src/types.tsstays at five factors.
The five lending factors
For each factor: the exact raw on-chain data that feeds it, the exact formula, and why the thresholds are what they are (anchored to an external/on-chain value where one exists, labeled an unvalidated judgment call where none does).
Two fixed-point scalars appear throughout, taken from
blend-contracts-v2/pool/src/constants.rs:
SCALAR_7 = 10^7— decimals forc_factor,l_factor,util,max_util.SCALAR_12 = 10^12— decimals ford_rate,b_rate.
A reserve's human-unit totals (used by several factors) are:
supplied = b_supply × b_rate / (SCALAR_12 × 10^assetDecimals)
borrowed = d_supply × d_rate / (SCALAR_12 × 10^assetDecimals)
1. collateralSafety — collateral concentration (weight 0.20)
What it measures: how spread out the pool's supplied value is across its reserves. A pool whose value sits in one asset is far more exposed to a single de-peg or liquidation cascade than a balanced one.
Raw on-chain data (Soroban RPC, no third party):
- Per reserve, from the pool contract's persistent storage:
ResDataentry →b_supply,b_rateResConfigentry →decimals
- Oracle price per asset:
lastprice(Asset::Stellar(address))on the pool's configured oracle contract →price, and the oracle'sdecimals(). - USD value per reserve:
suppliedUsd = supplied × (price / 10^oracleDecimals).
Formula — a normalized Herfindahl–Hirschman Index (HHI) over each reserve's share of total supplied USD:
Let vᵢ = supplied USD of reserve i (only priced reserves with vᵢ > 0)
n = number of such reserves
sᵢ = vᵢ / Σv (each reserve's share)
HHI = Σ sᵢ² (ranges from 1/n for a perfectly even split, to 1)
collateralSafety = clamp( (1 − HHI) / (1 − 1/n) × 100 , 0, 100 )
Edge cases: 0 priced reserves → 0 (can't assess, treated as unsafe rather than guessed); exactly 1 priced reserve → 0 (fully concentrated by definition).
Why HHI / why these anchors: HHI is the standard, widely-published concentration
measure (used by competition regulators and in portfolio analysis) — an external
framework rather than a Stenion invention. The anchoring points are not arbitrary either:
1/n (a perfectly even split) is the mathematically safest achievable state for n
reserves and maps to 100; 1 (everything in one asset) is the worst and maps to 0.
Normalizing by 1/n means the score grades a pool against the best it could do given how
many reserves it has, not against an arbitrary constant.
2. oracleSafety — price trustworthiness: freshness and manipulation resistance (weight 0.25)
Why this factor is not just price age. An age-only oracle factor scores a fresh but manipulated price 100 — which is precisely the configuration behind the February 2026 YieldBlox/Blend incident. Freshness alone is not a weak signal on that axis, it is a misleading one, so this factor takes the binding constraint of freshness and manipulation resistance. Earlier development-era scores did measure age alone; that history was discarded rather than carried forward, and no stored row was computed that way — see Current version.
What it measures: whether the prices this pool actually runs on can be trusted. Two things must both hold, and the factor takes the binding constraint of the two — a bounded stale price and a fresh unbounded price are both untrustworthy, for different reasons:
oracleSafety = min( priceFreshness , deviationBound )
Both sub-signals take the worst reserve, the same convention as every other factor:
the binding constraint is the single weakest reserve, and averaging would hide it. Both
are published in the factor's components array so the composite is never an opaque
number.
Both are also anchored to parameters the pool's own price path has to publish, which makes this the one factor with a precondition attached: a market whose oracle publishes neither a staleness tolerance nor a deviation bound is not scored at all, rather than scored with a guessed anchor or with this factor dropped. See 2e, the oracle-legibility precondition.
Every reserve at the binding value is named, not one of them. When several reserves tie
on a sub-signal the detail lists all of them; when all of them tie it says so rather than
singling one out. This is reporting only — the published value is the same minimum either way
— but it is load-bearing for reading a score honestly. Blend prices its whole pool from one
aggregator publish round, so its reserves carry identical ages and always tie on freshness.
Naming one of them would make an iteration-order artifact read as a diagnosis, and a reserve
name that is really a tie-break is worse than no name at all.
2a. priceFreshness — how stale the worst price is
Raw on-chain data (Soroban RPC): per reserve, the price's publish timestamp from
the method the protocol's own pool calls; fetchedAt is the adapter's read time;
age = fetchedAt − timestamp.
fresh = the protocol's own publish/refresh interval → 100
dead = min( protocol's own max acceptable price age, 3600s ) → 0
priceFreshness = clamp( (age − dead) / (fresh − dead) × 100 , 0, 100 )
No usable price for a reserve → 0 (a missing feed is maximally unsafe, not skipped).
Both anchors are the protocol's own on-chain parameters, the same anchoring pattern
utilizationSafety uses. Which parameter each resolves to is a documented per-protocol
fact, not a per-protocol rule:
| Protocol | fresh source | dead source |
|---|---|---|
| Blend | oracles()[i].resolution on the pool's oracle aggregator (300s) | max_age() on the same aggregator (900s) |
| Kinetic (K2) | PriceCacheTtl on the price oracle (30s) — the window inside which K2 itself treats a price as current; K2 exposes no publish-interval getter | the tighter of the per-asset max_age (43200s) and the global price_staleness_threshold (3600s) |
Taking the tighter of two limits a protocol declared is not a Stenion threshold — both numbers are K2's, and the binding one is the one that governs.
⚠️ The 3600s cap on
deadis the one Stenion constant left in this factor, and it is an unvalidated judgment call. Anchoring purely to a protocol's own max age would mean a protocol scores better for tolerating staler prices — K2's per-assetmax_ageis 12 hours, which would make a six-hour-old price score ~50. That is the wrong incentive for a platform protocols are ranked by, so the anchor is capped. There is no external framework fixing the cap at one hour; it is open to challenge like any threshold here. It lives in one place,STALE_CEILING_SECONDSincore/src/scoring.ts.
2b. deviationBound — can a single update move the price arbitrarily far?
Binary, not a curve:
deviationBound = 100 if the pool's price path bounds a single-step move, and that bound is armed
0 otherwise
| Protocol | Bounded when |
|---|---|
| Blend | the aggregator's per-asset max_dev satisfies 0 < max_dev < 100 — the contract's own condition in oracle-aggregator/src/price_data.rs |
| Kinetic (K2) | max_price_change_bps > 0 and get_last_price(asset) returns a present, non-zero baseline |
Why the extra clause for K2. The two contracts fail in opposite directions when there
is no prior price to compare against. Blend's aggregator fails closed: with no older
record it returns None and the reserve simply cannot be priced. K2's
validate_price_change fails open: with no stored baseline it returns Ok and lets
any price through, so a configured bound with no baseline is inert. Checking the baseline
is what distinguishes a breaker that is configured from one that is actually armed — and
the no-baseline case is exactly the newly-listed-thin-asset scenario that the YieldBlox
incident ran through.
Why this is anchored, and what it isn't. The scored quantity is the presence and
arming of a bound, and it is read from the protocol's own on-chain configuration — the
same pattern as utilizationSafety's max_util. max_dev = 0 does not mean "a tight
bound of zero"; the aggregator's own type documentation says "If this is 0, the oracle
will just fetch the last price within the resolution time" — the check is skipped
entirely. That is provably the condition that permits an unbounded single-step move.
Base assets are excluded, not scored 0. The Blend aggregator's lastprice
short-circuits its Base and BaseAssets to exactly 1.0 at the current ledger time
without consulting any upstream feed. Those reserves have no oracle-derived price to
grade, so they are dropped from both sub-signals and the count of excluded assets is
disclosed. (Whether such a peg holds is a real risk — but it is a collateral/peg
question, not an oracle-robustness one, and inventing a number for it here would be the
kind of fabrication ground rule 4 forbids.)
2c. Per-feed price ages are disclosed, never scored
priceFreshness grades the worst reserve, which is the right thing to score but hides the
spread — and on real data the spread is the informative part. A factor value of 0 reads as
a general condition of the oracle. "Two feeds have not updated in hours while two others update
every few seconds, through one contract and one source" is a specific, checkable statement
about which feeds are being maintained, and it is the one a depositor can act on.
So every reserve's price age is published as a disclosure-only component (priceAges,
value: null), ordered oldest-first, alongside a count of how many exceed the protocol's own
declared staleness limit — Blend's aggregator max_age, K2's price_staleness_threshold.
The count is therefore a statement about the protocol's own rules, not about a Stenion line. A
reserve with no usable price at all sorts as the oldest rather than the freshest.
It is not scored, because it would double-count: these are the same ages priceFreshness was
computed from, republished so that the grading can be checked rather than taken on faith. It is
published on healthy pools as well as unhealthy ones — a disclosure that appears only where
trouble is expected gives a reader no baseline to compare against.
2d. Bound tightness is disclosed, never scored
The raw bound is published as a disclosure-only component (value: null) — visible,
never graded. Grading it would invent comparability the underlying data does not support:
Blend max_dev | K2 max_price_change_bps | |
|---|---|---|
| Scope | per asset | global |
| Units | whole percent (60 = 60%) | basis points (2000 = 20%) |
| Baseline | the previous upstream record, one resolution step back | get_last_price — the last price the contract served |
| Bounds move per | publish interval (300s) | query — no fixed time spacing |
Both compute |new − old| / old, and the unit difference normalizes trivially. The
baseline difference does not: "20% per arbitrary interval" and "60% per five minutes" are
different quantities, so the intuitive reading that K2's bound is three times tighter than
Blend's is unsound. Publishing the numbers side by side without a score is the honest
treatment.
2e. The oracle-legibility precondition
Both halves of this factor are anchored to parameters the pool's own price path
publishes: §2a's window comes from resolution and max_age, §2b's bound from per-asset
max_dev. That is the whole design — the numbers are the protocol's, not Stenion's. It has
a precondition hiding inside it, which this section makes explicit:
A market is scorable only if its price path publishes the parameters §2 grades against. Where it does not, the market is not scored at all — it is not scored with a guessed anchor, and it is not scored with
oracleSafetyomitted.
This is the market-size floor one factor down, and it has the same
shape: a precondition on what gets scored at all, rather than a rule about how to score it.
Markets excluded by it are published in
dashboard/app/lib/coverage.ts as oracle-not-gradable,
with a per-market reason and a date — never as a protocols row, and never with a numeral.
Where the line falls today. Blend's oracle-aggregator publishes all three reads
(max_age(), oracles(), asset_configs()); those are Blend's interface, not SEP-40's.
Every Blend V2 pool Stenion scores sits on one. Four live pools do not, and the interfaces
below were read on 2026-08-26 out of each oracle's own wasm (contractspecv0 via Soroban
RPC getContractMethods) rather than probed by calling a list of guessed names — a guess-list
cannot distinguish "this contract lacks the method" from "this is a different contract":
| Pool | Oracle wasm | What the contract actually is | max_age | oracles | asset_configs |
|---|---|---|---|---|---|
| Blend Fixed | 41df0489… | Blend oracle-aggregator | ✅ | ✅ | ✅ |
| YieldBlox | 8cf43882… | oracle-aggregator, different build | ✅ | ✅ | ✅ |
| Etherfuse | 65300c00… | oracle-aggregator | ✅ | ✅ | ✅ |
| Orbit | a71a844e… | bridge oracle — ctor (admin, stellar_oracle, other_oracle) | ~ | ~ | ~ |
| Forex | 1d1c90d3… | proxy — CONFIG.base_oracle points one hop up at a SEP-40 feed | ~ | ~ | ~ |
| Spectra PTs | 4a444181… | deterministic zero-coupon-bond pricer — not a feed at all | ~ | ~ | ~ |
| Solv | 5700be21… | SEP-40 feed registry — the only SEP-40 implementer of the four | ~ | ~ | ~ |
Those four are not one shape. They are four different contracts with four different wasm
hashes doing four different things, and the only thing they agree on is the column that
decides this: none answers any of the three. Worth stating because the obvious fix — "handle
the other oracle shape too" — is really "handle four more shapes", and each would be a
separate reading of a separate contract's semantics. All four answer decimals() and
lastprice(), so §1, §3, §4 and §5 compute normally for them; what is missing is only the
metadata §2 grades against.
Why there is no fallback anchor: SEP-40 does not define one
Stated plainly rather than worked around. SEP-40
defines base, assets, decimals, resolution, price, prices, lastprice — no maximum
acceptable price age and no deviation bound anywhere in the interface. The spec puts
staleness checking on the consumer ("Always check retrieved price data for staleness by
comparing the quoted timestamp with current date"), which is precisely the judgment §2 exists
to make and precisely what it refuses to make from an invented number.
So the only candidate is resolution() — a publish interval, not a staleness tolerance —
and using it fabricates a 100 on a demonstrably stale price:
Solv publishes
resolution() = 43200(12 hours, and mutable after deployment via its ownset_resolution, which its source documents as a deliberate deviation from SEP-40). Fed tofreshnessWindowwith nomax_ageto pair it with, that yields{fresh: 43200, dead: 86400}— becauseSTALE_CEILING_SECONDSclampsdeadand notfresh, so a feed that declares a slow tick gets a slow dead line rather than a capped one. Solv's genuinely stale feeds, read at 10,285s and 21,739s old on 2026-08-26, would both publishpriceFreshness100.
And resolution() is only present on one of the four at all. Orbit and Forex expose it one
hop upstream through their bridge/proxy, which would make the anchor a property of a contract
the pool does not itself publish; Spectra has no upstream feed to chase.
Two of the four price off the ledger clock, so freshness is 100 by construction
The sharper problem, and the reason this is a precondition rather than a "weak signal":
- Spectra PTs runs "Spectra Deterministic Oracle — Zero Coupon Bond Model". Its price is a
function of
start_t,maturityandinitial_implied_apyevaluated at the current ledger time, itslastpriceignores the asset argument entirely (its own doc comment says so), and its owner can move the target withset_future_pt_value. There is no publish event, so there is no such thing as a stale price:ageis always ~0 and any freshness formula returns 100 permanently, whatever happens to the asset. - Orbit's dominant reserve — 99.5% of that pool's $190,863 on 2026-08-26 — returns
exactly
1.0at current ledger time while touching no upstream contract. §2b already handles that case on an aggregator: base assets are excluded rather than scored, because there is no oracle-derived price to grade. Orbit's bridge publishes nobase(), so there is nothing to detect them with, and the reserve would be graded as a fresh feed.
A fabricated 100 is worse than a fabricated 0, because 0 at least renders in the danger band where a reader discounts it. Ground rule 4 forbids both.
What was rejected, and why it must not be re-proposed
-
Make the three calls optional and score
oracleSafetyon what remains. Rejected: there is nothing to score on. Both anchors are gone,deviationBoundcollapses to a constant 0 for all four pools — a constant is not a measurement — andpriceFreshnesscollapses to a constant 100 for two of them. It also reads as a diagnosis it did not make: amax_devof 0 means a protocol disabled its bound, which is the YieldBlox finding; publishing the same 0 for a protocol whose oracle never had the mechanism asserts a choice nobody made. -
A
nulloracleSafety, scored on the remaining four factors. Rejected, and this is the one that looks reasonable until it is measured.scoreFactorsrenormalizes over non-null weights, so droppingoracleSafetydivides by 0.75 instead of 1.00. Run against live chain data on 2026-08-26, that is not neutral — it is a large upward revision, becauseoracleSafetyis the heaviest factor (0.25) and the one a badly-configured pool scores worst on:Pool Score with oracleSafety: nullFor comparison, live registry Orbit 71 Blend 51Spectra PTs 49 Kinetic 27Solv 15 YieldBlox 25Forex 11 Orbit would publish 71 — the highest number in the registry, twenty points above Blend Fixed — while Admin-Frozen, 99.5% concentrated in a synthetic priced at a hardcoded 1.0, through an oracle publishing neither a staleness tolerance nor a deviation bound. That is an incentive inversion: YieldBlox scores
oracleSafety0 for having a deviation bound and disabling it, so not publishing the mechanism at all would pay better than publishing it switched off. Ship an opaque oracle, get a better number — which is ground rule 2's concern arriving through the back door.It is not the same failure as the rejected sixth factor under Factor weights: that one moved every protocol's score by changing the denominator for everyone. This moves only the affected pool's. The failure here is different and, in a ranked list, worse — two entries in the same ranked column would be graded on different rulebooks, four factors against five, which ground rule 1 does not permit. A
nullfactor is defined as "genuinely doesn't apply to this protocol" (the dashboard renders it "Not applicable to this protocol"). A pool that runs on prices, whose price configuration we could not read, is not a pool to which price trustworthiness does not apply. -
Reading the anchor from the upstream oracle a bridge or proxy forwards to. Rejected: it answers for a contract the pool does not publish and does not itself constrain, it exists for only two of the four, and following it is a bespoke traversal per oracle implementation — a per-market rulebook in all but name. It also would not have helped: the upstream feeds behind Orbit and Forex publish
resolution()and nomax_age, so the traversal lands back on the fabricated anchor above.
This does not bump lending's methodology version. It moves no published number and changes no
formula. Every market Stenion scores runs on an aggregator, so oracleSafety is computed
byte-identically before and after; what changed is that a precondition already implicit in §2
is now written down and enforced in code
(ORACLE_GRADING_READS and oracleNotGradable), instead of surfacing as
an unexplained HostError from whichever read happened to run first. Same reasoning, and the
same conclusion, as the market-size floor.
⚠️ What this costs, stated rather than glossed. Four live markets stay unranked, one of them holding real money — Orbit's $190,863 on 2026-08-26. The market-size floor's warning applies here in full: an excluded market has no entry at all, which is a stronger action than excluding a reserve.
coverage.tsis the answer to that and not a formality — each of the four gets a page, a per-market reason, the contract addresses, and averifysentence, so a reader can disagree with the decision from the same data it was made on. What they do not get is a number, because there is no number to give them.The precondition is a property of the oracle, not a verdict on the protocol. Nothing here says these markets are unsafe, or that their oracles are bad ones. A deterministic bond pricer is a perfectly coherent way to price a principal token; it is simply not a thing this factor knows how to grade. If such an oracle later publishes a staleness tolerance and a deviation bound, the pool becomes scorable with no rule change — it is a
BLEND_POOLSentry and a deleted coverage entry, in one PR.
What was considered and deliberately rejected
Recorded so these are not re-proposed as improvements later. Each was investigated against the February 2026 YieldBlox incident — the test being whether it would have distinguished the manipulated price from a legitimate one at the time, since a signal that looks sophisticated but would not have caught the actual attack is worse than none: it manufactures confidence.
-
Filtering
oracleSafetyby reserve size, the way §4/§5 are filtered. Rejected on principle, not on impact — and the distinction it turns on is the reason §4/§5 may be size-filtered while this factor may not:§4 and §5 measure current state. §2 measures a vulnerability. How drained a reserve is right now means little when the reserve holds $4, because the exposure is capped by what is actually in there. Whether a price can be trusted is not capped that way, because the attacker's move is to grow a position against the mispriced asset. A dust reserve with a stale price is an open door, not a small room. Its balance today says nothing about what can be borrowed against it tomorrow.
And it would blind the factor to the exact scenario it exists for: a newly-listed thin asset with a bad price is the shape the February 2026 YieldBlox incident ran through, which §2b already names. A filter that removes thin assets from an oracle-trust factor removes the attack it was built to catch.
-
A Stenion-computed deviation from the oracle's price history (calling Reflector's
prices(asset, N)ourselves and comparing the latest price to a trailing mean). Rejected — this would have made the platform actively worse. It is a coincident indicator, not a leading one: it can only fire while an attack is in progress, and only if the indexer happens to sample inside the manipulation window. The indexer runs every five minutes, so the overwhelmingly likely outcome is that it reads clean and Stenion publishes a confidentoracleSafetyof 100 during an active exploit. It also measures a code path the pools never consult: neither Blend's aggregator nor K2's oracle exposes price history to the pool at all. A signal that is usually silent during the event it claims to detect, computed over data the protocol does not use, is not a weak signal — it is a misleading one. -
TWAP. Not available: the deployed Reflector contracts (
version() == 6) expose notwapmethod — the exported interface isbase, assets, decimals, resolution, price, prices, lastprice, last_timestamp, history_retention_period, …. Earlier Reflector versions had one; the live contracts do not. Neither protocol's oracle passes history through either. And on the merits it would not have helped: the attacker held the only trades in the window, so a short TWAP over a dead order book is the manipulated price. -
Oracle type / provider identity ("is it Reflector?"). Zero discriminating power: the exploited pool and the healthy Blend pool both price through Reflector-family feeds via the same
oracle-aggregatorcontract family. What differed was configuration, not provider. -
Number of upstream sources. Both pools had exactly one upstream oracle, so it would not have separated them. It is also not comparable across protocols: a count is only readable where a contract happens to publish one (K2's upstream RedStone adapter exposes
unique_signer_threshold() == 3; Reflector's node consensus is not exposed on-chain at all), so counting would systematically understate feeds that keep their aggregation internal. -
SDEX order-book depth via Horizon. Conceptually the right quantity — thin market depth is what made the manipulation cheap — and mechanically readable. Rejected for now on three grounds: order books are trivially spoofable with walls that are never hit; it only applies to assets priced off the Stellar DEX; and it cannot be validated retroactively, because the exploited market has since been rebuilt. Tracked as a candidate in
ROADMAP.mdrather than shipped on intuition.
What this factor would have said on 2026-02-22
Running the shipped adapter against the exploited pool and the healthy one today, same rulebook, no special-casing:
| Pool | priceFreshness | deviationBound | oracleSafety |
|---|---|---|---|
Blend Fixed V2 (CAJJZSGM…) | 100 | 100 — all reserves bounded (max_dev 60/20/20) | 100 |
YieldBlox (CCCCIQSD…, the exploited pool) | 84 | 0 — XLM and AQUA carry max_dev: 0, check disabled | 0 |
Both pools' prices are fresh, so an age-only factor scores both high — 100 and 84, the latter being an ordinary mid-window price age, not a warning. This factor separates them anyway, and on the axis that actually failed.
This is no longer a demonstration run. As of the multi-pool change, the YieldBlox pool is a registered, continuously scored entry in the public registry, and the row above is its live
oracleSafety, published every five minutes like any other. Two consequences worth stating: the claim in this section is now checkable by anyone againstGET /api/v1/protocol/yieldbloxrather than reproducible only by running the adapter by hand; and the number will move, because it is live. The pairing that matters — a fresh price and a disabled bound — is a property of the pool's configuration, not of the moment it was sampled.The entry is labelled a Blend V2 pool wherever it appears (
deployedOnon both API responses). It is not a third protocol, and the registry must not be read as saying so.
⚠️ Two honest limits on that claim, stated rather than glossed:
- The historical
max_devis a deduction, not a reading. Soroban RPC serves no historical contract state, so the exact value USTRY carried on 2026-02-22 cannot be read back. What is verifiable: the deployed aggregator skips the deviation check entirely whenmax_devis0or≥ 100, and rejects the price outright otherwise — so a ~100× single-step move is arithmetically incapable of passing any bound between 1 and 99. USTRY's bound must therefore have been disabled. USTRY today carriesmax_dev: 10; XLM and AQUA in that same live contract still carry0.- Semantics were verified against the public repo, not that binary. The exploited pool's aggregator (wasm
8cf43882…) and Blend Fixed V2's (41df0489…) are different builds. Both export the same eleven functions, and themax_devlogic above is read from blend-capital/oracle-aggregator; it has not been decompiled from the exploited pool's specific binary.
⚠️ K2's enforcement is an inference, held to the same standard.
max_price_change_bpsis enforced on every return path ofget_asset_price_datain K2's audited source (code-423n4/2026-04-k2), and the deployed wasm contains both themax_price_change_bpsandPriceChangeTooLargesymbols. But the livekinetic_routerdoes not call that method — it callsget_asset_prices_vec_fresh, one of nine functions present in the deployed oracle and absent from the audited source, whose source is not public. The audited siblingget_asset_prices_vecdoes enforce the breaker, and all three methods return identical data today. We score it as enforced on that basis. That is an inference, not a verification, and it is written up as a finding in its own right — see the Kinetic entry in the registry.
3. adminKeySafety — admin signer structure + activity (weight 0.20)
What it measures: how much unilateral, live control a single party has over the pool. A lone hot key that can reconfigure the pool is the sharpest centralization risk; multisig and inactivity are safer.
Raw on-chain data:
- The admin address comes from the pool contract's instance storage (
Admin, orConfig.admin) via Soroban RPC. - If the admin is a keypair account (
G…), signer structure and activity come from Horizon (official Stellar infra, not a third party):GET /accounts/{address}→thresholds.high_threshold,signers[](→signerCount)GET /accounts/{address}/operations?order=desc&limit=200→created_atof each op;recentOps= count within the last 30 days.
- If the admin is a contract (
C…), Horizon has no account entry to introspect — there is genuinely nothing to measure.
Formula — a tiered base (categorical, NOT a curve) minus a continuous activity penalty:
This factor is deliberately tiered, not a continuous function, because signer structure is categorical. The base value is chosen by tier:
| Tier | Base | Detected by |
|---|---|---|
| Contract-governed admin | 60 | admin address starts with C… (flagged neutral baseline — see below) |
| Single master key | 40 | keypair account, not multisig |
| N-of-M multisig (N ≥ 2) | 90 | signerCount > 1 AND high_threshold > 1 |
| Multisig + timelock | 100 | RESERVED — see note |
Then a continuous activity penalty is subtracted:
activityPenalty = min(30, recentOps × 3) # capped so structure still dominates
adminKeySafety = clamp( base − activityPenalty , 0, 100 )
⚠️ The "Multisig + timelock" (100) tier is reserved and not yet reachable. No on-chain timelock signal is exposed to the adapter through Horizon today, so nothing is ever scored 100 by this factor at present. It is documented as the intended top tier so that when a timelock signal becomes detectable, the tier already exists rather than being invented ad hoc. This is an aspirational placeholder, explicitly flagged, not a live rule.
⚠️ The contract-governed baseline (60) is a flagged neutral value, not a measurement.
When the admin is a contract, we cannot introspect its governance via Horizon. Rather than
fabricate a plausible signer/activity number, we assign a fixed, clearly-labeled neutral
baseline of 60 and say so in the factor's detail string. This is honest ignorance, not a
score.
Why these numbers (unvalidated judgment calls, partially anchored):
- The single-key (40) vs multisig (90) split is anchored to a real, hard security fact: a
1-of-1 key is a single point of unilateral compromise; an N-of-M multisig with
high_threshold > 1provably requires more than one party to reconfigure the pool. The detection condition (signerCount > 1 AND high_threshold > 1) reads Stellar's actual account threshold model, not a proxy. - The exact base values (40, 90, 60) and the activity penalty shape (
−3per op, capped at−30) are unvalidated judgment calls. The cap deliberately keeps structure dominant over activity (a busy multisig should still beat an idle single key). There is no external framework these specific integers are anchored to — they are open to challenge.
4. liquiditySafety — free-liquidity depth (weight 0.15)
What it measures: the absolute withdrawal/liquidation cushion — how much value could
leave before the pool is drained. Distinct from utilizationSafety, which measures
proximity to the configured cap rather than absolute headroom.
Raw on-chain data (Soroban RPC): per reserve, ResData (b_supply, b_rate,
d_supply, d_rate) and ResConfig (decimals), used to compute supplied and
borrowed per the totals formula above.
Formula — free-liquidity share of the worst reserve, over reserves large enough to be scored (see The minimum-size filter below):
For each reserve with supplied > 0 that passes the minimum-size filter:
free = clamp( (supplied − borrowed) / supplied × 100 , 0, 100 )
liquiditySafety = min(free) across all such reserves # worst reserve wins
Edge cases, both → 0: no reserve with supplied > 0, and every reserve excluded by
the minimum-size filter. A minimum over an empty set is undefined, not the top of the scale
— an unassessable pool is reported as unassessable, the same way §1 treats having nothing to
price. Returning 100 here would publish "maximally safe" derived from no data, which ground
rule 4 forbids. The two are reported with different detail strings: "the pool is empty" and
"everything in it is too small to grade" are different findings.
The minimum-size filter
Applies to liquiditySafety (§4) and utilizationSafety (§5) only, identically for every
protocol.
The problem it solves. Both factors select the worst reserve, so a reserve holding
effectively nothing can set a protocol's published number. On the 2026-08-16 Kinetic snapshot a
$3.00 PYUSD reserve — 0.19% of a $1,571 pool — was the worst reserve on both factors and
set liquiditySafety to 34 and utilizationSafety to 18. Nobody's capital was meaningfully
exposed to it. That is a misleading number, not a conservative one.
The rule. A reserve is scored if either test passes, and excluded only when both fail:
| Leg | Test | Anchor |
|---|---|---|
| A | suppliedUsd ≥ the protocol's own declared minimum viable exposure | the protocol's own on-chain parameter, where it declares one |
| B | suppliedUsd ≥ 0.5% of the pool's own total supplied USD | none — an unvalidated judgment call (see below) |
Leg A is per-protocol in exactly the sense §5's cap is: the pattern ("grade against a
parameter the protocol set itself") is the invariant, and which parameter it resolves to is a
documented per-protocol fact.
| Protocol / market | Leg A source | Value |
|---|---|---|
| Blend — Fixed V2 | PoolConfig.min_collateral, read live from pool instance storage, denominated in the oracle's base asset (Other:USD, 7 decimals) | 50000000 = $5.00 |
| Blend — YieldBlox V2 | the same field, read live from this pool's own instance storage — read per pool, never inherited from the flagship | 50000000 = $5.00 |
| Kinetic (K2) | none — K2 declares no minimum-exposure parameter on chain. Leg B alone applies. | n/a |
Both live Blend pools happen to declare the same floor. That is a coincidence of their
configuration, not a property of the adapter: leg A is resolved from whichever pool an
adapter instance was pointed at, and a Blend pool declaring a different min_collateral
would be graded against its own.
min_collateral is Blend's own dust guard: the smallest collateral a position may hold and
still borrow, set where liquidating a position stops being economically worthwhile. A reserve
whose entire supplied value sits below it cannot host even one position the protocol itself
considers viable. That is the same question this filter asks, which is why it is borrowed
rather than invented.
K2's absence is verified, not assumed. The router's instance storage and every reserve's
ReserveConfiguration bitmap were read looking for an equivalent. What K2 exposes is MINSWAP
(a slippage bound), FLPREMMAX, HFLIQTH/PLIQHF (health-factor lines) and a supply/borrow
cap pair in data_high — all maxima or unrelated. If K2 ever ships a minimum, leg A turns on
for it with no rule change.
Why both legs, and not one. Each covers a failure the other has, both demonstrated on live data:
- Absolute-only breaks a small pool. Any floor sized for a real market ($1k, $10k) excludes all four of K2's reserves — its entire pool is ~$1,500. Both factors would go to cannot-assess and K2's score would drop, from 28 to 15. Worse than the problem.
- Relative-only breaks a large pool. 0.5% of Blend's $186M is ~$928,000, so a reserve holding half a million dollars of real capital would be silently dropped. Leg A keeps it at $5.
The 0.5% in leg B is an unvalidated judgment call. There is no external or on-chain framework fixing it; with
STALE_CEILING_SECONDS(§2) it is one of only two Stenion-chosen constants left in the continuous factors, and it is open to challenge like any threshold here.It is deliberately set at the low end of the band that works, because the two directions of error are not symmetric. Too low leaves a dust reserve in, which reports a misleading number. Too high excludes a small but genuinely-used reserve, which hides real risk — strictly worse. 0.25% would have flipped on the live K2 reserve between two consecutive days ($3.00, then $4.00, against a $3.85 line); 0.5% clears it both times with margin.
Excluded reserves are disclosed, never silently dropped. Each affected factor publishes an
excludedReserves component with a null value — the same "measured, shown, deliberately not
graded" form as §2c/§2d — naming each excluded reserve, its supplied USD, its share of the pool,
and the score it would have contributed. A reader can therefore see the number the filter
suppressed and disagree with the exclusion, instead of never learning of it.
⚠️ This filter gives §4 and §5 an oracle dependency they did not previously have. Both are otherwise pure balance ratios that need no price at all; the filter is USD-denominated. When no reserve can be priced, the filter does not run and every reserve is scored — the two factors degrade to exactly their pre-filter behaviour rather than refusing to score. That is the right fallback, but it means a pool's liquidity and utilization numbers mean something slightly different during an oracle outage: they are unfiltered, and a dust reserve can bind them again. An individual unpriced reserve is likewise kept, never read as worthless — "could not measure" is not "empty".
The filter cannot empty the scored set on a real pool. Shares sum to 1, so the largest
reserve always holds at least 1/n, which clears 0.5% for any n ≤ 200. The all-excluded
branch above is therefore unreachable in practice — it is implemented and tested synthetically
anyway, because that is precisely where a "cannot assess" could quietly become a 100 again.
⚠️ OPEN QUESTION, raised by the YieldBlox pool and deliberately not resolved here. The two legs are OR'd, so leg A can override leg B — and on a small Blend pool it overrides it almost entirely. YieldBlox holds ~$1.28M, putting leg B's 0.5% line at ~$6,396; six of its eight reserves fall below that line ($39.47 to $4,243.73) and every one is scored anyway, because Blend's $5
min_collateralpasses for all of them. The result is thatliquiditySafety(10) andutilizationSafety(0) are both set by a reserve holding $1,096.85 — 0.086% of the pool.That is the shape of the problem this filter was added for. On Blend's Fixed pool it is invisible: leg A is a documented no-op there, because the smallest reserve holds $3.4M. On a pool three orders of magnitude smaller, the same $5 floor is doing all the work and leg B's guard never engages.
It is recorded, not fixed. Changing it — sizing leg A relative to the pool, capping it, or making the legs AND rather than OR below some pool size — moves published numbers on a live entry, and is a threshold change under the same review bar as any other (see Disputing or changing a threshold). It is equally arguable that the current behaviour is correct:
min_collateralis the pool's own statement of the smallest position worth liquidating, and a $1,097 reserve at 90% utilization is a real reserve with real depositors, not the $3.00 dust the filter was built to exclude. What is not defensible is leaving the tension undocumented, which is why it is written down here.
Why this shape / this anchor: (supplied − borrowed) / supplied is 1 − utilization,
i.e. the fraction of supplied value that is actually withdrawable right now. That is a
direct on-chain quantity, not a modeled one — the anchor is the pool's own balances. Taking
the worst reserve rather than a pool-wide average is deliberate: liquidity crises happen
in the single most-drained reserve, and averaging would hide it. The mapping (free % → score
%) is 1:1 and intentionally has no free parameters to tune, so there is nothing arbitrary to
anchor.
The market-size floor
The minimum-size filter one level up. That filter asks whether a reserve is big enough for its number to mean anything; this asks the same of a whole market. They are two halves of one idea — a size below which a published number stops carrying information — and they are written together so neither looks like an afterthought.
The problem it solves. K2 deploys its markets as separate router contracts running identical code, the same way Blend's factory deploys pools. Three are live on mainnet as of 2026-08-20, and two of them are empty:
| Market | Reserves | Total priced supplied value |
|---|---|---|
K2 primary (CCTUJZLY…) | USDC, XLM, PYUSD, SolvBTC | $1,781 |
K2 SolvBTC/xSolvBTC iso (CCGXGXIL…) | SolvBTC, xSolvBTC | $3.62 |
K2 Earn / earnUSDC (CDWPVHKB…) | USDC, earnUSDC | $0.00 |
Point the shipped rulebook at either of the bottom two and it does not fail — it returns a score. Every factor falls to its can't-assess branch, and every one of those branches is 0. So a market holding nothing publishes 0, in the danger band, reading "this is dangerous" to anyone scanning the registry when what is true is "there is nothing in here." That is the same misleading-number failure §4 and §5 were filtered for, one level up, and it is worse at this level: a filtered reserve still leaves a scored market with a disclosure beside it, whereas this is the market's entire published number.
The rule. A market is scorable only if it can hold at least one position the protocol itself considers viable:
| Leg | Test | Anchor |
|---|---|---|
| A | total priced supplied USD ≥ the protocol's own declared minimum viable position | the protocol's own on-chain parameter, where it declares one |
| B | none — no relative leg exists at this scale. See below; this is a real gap, not an omission | — |
| Protocol / market | Leg A source | Value |
|---|---|---|
| Blend (both pools) | PoolConfig.min_collateral, read per pool | 50000000 = $5.00 |
| Kinetic (K2) | none declared on chain — the $5.00 above is borrowed as an analogue, and is a flagged judgment call for K2, not an anchor | $5.00 |
This is the same parameter, and the same reasoning, that §4/§5's leg A already uses:
min_collateral is the protocol's own statement of the smallest collateral a position may
hold and still borrow. A market whose entire supplied value sits below it cannot host even
one position the protocol itself would let borrow. There is nothing there to assess, and the
number is not a measurement of risk — it is a measurement of absence.
Against the table above: K2 Earn fails on total supplied value of exactly zero. The SolvBTC/xSolvBTC market fails at $3.62. K2's primary market clears by three orders of magnitude, as do both Blend pools.
Why there is no relative leg, unlike §4/§5. The reserve filter has two legs because each covers a failure the other has. No such second leg exists here. Relative to the market's own reserves is what §4 and §5 already do. Relative to the other markets in the registry would make one market's listing depend on another market's size — a market could become unlistable because a different one grew, while nothing about its own on-chain state changed. A rule about a market's own data must not have that property. So this floor is absolute-only, which is precisely the shape §4/§5 rejected as insufficient on its own, and that limitation is the reason it is set low rather than at a number that sounds meaningful.
The direction of error is deliberately toward keeping markets in. Two reasons, and the asymmetry is not the same one §4/§5 reasoned about:
- Raising it buys nothing against the failure it exists for. A score computed from no data is fully prevented at $5. Every dollar above that excludes markets that genuinely can be assessed, in exchange for nothing.
- Excluding a market is a much stronger action than excluding a reserve. A filtered reserve leaves a scored market and a published disclosure naming what was suppressed. An excluded market has no entry at all — no score, no factors, no disclosure, nothing for a reader to disagree with. §4/§5 already call hiding a small-but-real reserve "strictly worse" than leaving a dust one in; at market scale that error hides everything at once.
What this floor guarantees — and what it does not. It guarantees only that a published number was computed from something rather than nothing. It is emphatically not a quality bar: a market holding $50 clears it, and its score would still be close to meaningless. Saying so plainly matters more than the threshold does, because a floor that sounds like a meaningfulness test while being a scorability test is worse than no floor.
⚠️ A separate question this deliberately does NOT answer: is a scorable market worth listing? K2's primary market is a registered, ranked entry holding $1,781. It clears this floor by three orders of magnitude and is still small enough that a reasonable person could ask whether ranking it beside a $185M pool conveys what the ranking appears to convey.
That is a curation question — what belongs in the registry — not a question about whether a number can be computed, and answering it with a threshold in this document would dress an editorial judgment as a measurement. It also has a consequence a scoring threshold does not: any such bar set above $1,781 would delist a live entry, breaking a public URL (
/protocol/kinetic) and orphaning a published history. Flagged here, resolved nowhere yet.
⚠️ An excluded market has NO score. It does not have a score of zero. This is the market-level form of the warning under §4/§5 about "cannot assess" quietly becoming a number again, and it is the whole reason the floor is written down.
- It must mean: the market is not registered. If a registered market later falls below the floor, its published
safetyScoremust becomenull— the never-scored representation the API already defines and the dashboard already renders as an em dash.- It must never mean: a score of 0 (which renders in the danger band and says the opposite of what is true), a score of 100, or — the live hazard — registering the market and letting the five factors fall to their can't-assess branches, which is exactly what the shipped code does today and exactly how an empty market publishes a 0.
Enforcement, stated honestly: this floor is currently enforced only by the decision not to register such a market. No code path implements it. Nothing in an adapter, the indexer or the store can express "this market is not scorable" as distinct from "this market scored 0", so a registered market that drained below the floor would keep publishing a number today. Closing that needs a distinct not-scorable outcome through
AdapterandRunRecord; it is filed inROADMAP.mdrather than implied to exist here.
This does not bump lending's methodology version. It moves no published number: every market Stenion currently scores — both Blend pools and K2's primary market — clears the floor, so no stored score is computed differently and none becomes non-comparable. It documents a precondition on what gets scored at all, which is additive.
5. utilizationSafety — headroom below the configured cap (weight 0.20)
What it measures: how close live utilization is to the protocol's own on-chain utilization stress line — the point the protocol itself defines as "borrowing should stop growing here." Approaching it is a concrete, protocol-defined stress signal.
Formula — headroom below the protocol's utilization line, worst reserve:
For each reserve with supplied > 0 and cap > 0 that passes the minimum-size filter:
util = borrowed / supplied # computed LIVE from balances, not a config field
headroom = clamp( (cap − util) / cap × 100 , 0, 100 )
utilizationSafety = min(headroom) across all such reserves # worst reserve wins
The minimum-size filter is §4's, unchanged and applied identically here — one rule, both factors.
Edge cases, all → 0, for the same reason as §4: no reserve with supplied > 0, every
reserve excluded by the minimum-size filter, and no reserve with cap > 0. The second is the sharper one — reserves can hold real debt while
declaring no utilization ceiling at all, and grading that as full headroom would measure distance
to a line nobody set. The two are reported with different detail strings, since "the pool is
empty" and "the pool declares no ceiling" are different findings.
cap is per-protocol — it is always the protocol's own on-chain utilization parameter,
never a Stenion constant. Which parameter that resolves to:
| Protocol | cap source | Meaning of the line |
|---|---|---|
| Blend | per-reserve max_util (ResData/ResConfig, 7-dec fixed point → max_util / SCALAR_7) | a hard throttle — Blend throttles and eventually pauses borrowing as utilization nears max_util |
| Kinetic (K2) | OPTIMAL_UTILIZATION_RATE = 0.80 (contracts/shared/src/constants.rs) | the interest-rate kink — past 80% util, K2's Aave-V3 rate curve steepens sharply to discourage further borrowing |
Why this anchor (the strongest in the set): the threshold is not a Stenion constant at all — it is the protocol's own on-chain parameter. The formula grades each reserve against the exact line the protocol configured, so the "danger line" is set by the protocol, not by us. This is the pattern every continuous factor should aspire to. Worst-reserve selection is deliberate for the same reason as liquidity — the binding constraint is the single reserve closest to its line.
⚠️ Two honest caveats on the K2 anchor (flagged, not hidden):
- K2's kink is a rate inflection, not a hard pause — past 80% util K2 keeps lending (just expensively), whereas Blend's
max_utilis an actual throttle. The two lines mean slightly different things; the formula treats "distance to the protocol's declared utilization ceiling" uniformly, which is the intended abstraction.OPTIMAL_UTILIZATION_RATEis read as K2's global default (0.80). Per-reserve kink overrides, if any, live in K2'sinterest_ratestrategy contract, which is out of scope in the audited source (code-423n4/2026-04-k2) and so not independently verifiable — if a reserve overrides the default this factor uses the documented 80%, not that reserve's exact kink. Revisit if K2 exposes a readable per-reserve optimal-util.
Dex
Everything from here to the end of this file is the DEX rulebook and the DEX rulebook alone —
its version changelog, its two factors, what it refuses to score, and what it defers. Nothing in index.md is lending-specific and all of it applies here; nothing
in lending.md may be assumed to hold for this category, and nothing here may be
assumed to hold for lending.
Which markets are scored under it: one — Aquarius's XLM/USDC constant-product pool
(CA6PUJLB…, registry id aquarius-xlm-usdc). This section was published and reviewable before
any market was scored under it, which ../TAXONOMY.md says is the point of
writing it first. It was admitted as a gate-checked submission against Aquarius on Stellar mainnet
in two reviews — the factor set first, the weight table second — then implemented in
adapters/aquarius/ and registered.
One market, and the reason is the indexer rather than the rulebook. Aquarius runs 340 pools
across 304 token sets — read from the router's own get_pools_for_tokens_range at ledger
64,182,824 on 2026-08-29 — and every one of them is scorable under the rules below, because neither
factor here is size-sensitive (see Size floor: none, and none pending).
What limits the registry is that one scoring cycle runs inside a 60-second serverless ceiling and
fits five markets, of which four were already lending. The other 339 are published as assessed and
unregistered on the registry, under coverage.ts's awaiting-capacity status — never as a score,
and never as a finding about them.
This rulebook is complete: two factors, each with a formula, and a reviewed weight table. The table was deliberately absent when the category was admitted and landed in a review of its own — a weight can only ever be a type-(b) unvalidated judgment call, and arguing one alongside the argument for the factor set would have buried it. Both halves are here now, in Factor weights and The two dex factors, and
CATEGORY_FACTORS.dexin../core/src/weights.tscarriesstatus: 'published'to match. Every Stenion-chosen number in either formula is listed in Unvalidated judgment calls.The category ships with TWO factors, not the three the submission proposed.
depthSafetyis deferred by question A, which is resolved rather than open. Gate 0 is re-argued for the two that remain in Gate 0, re-argued.There is no size floor, and none is pending. Neither surviving factor is size-sensitive, so there is nothing for a floor to protect — see Size floor: none, and none pending. This is a consequence of question A's resolution, not an open item.
Everything below was read from mainnet, not from documentation. That distinction is
load-bearing for this category in a way it was not for lending: github.com/AquaToken/soroban-amm
— the repository Aquarius's own audit scope links to — returns 404 as of 2026-08-27. There is
no source to read. Every interface claim here comes from the contract spec in the deployed wasm
(getContractMethods) and from instance storage, which is what Gate 8 asks for anyway.
Two dated read sets appear below, and each reading says which it is from. The census
readings — the 340-pool survey, the role structure, the issuer-flag sample, the kill-switch state —
are mainnet ledger 64,152,946, 2026-08-27T20:00Z, taken when the category was admitted. The
worked example is computed from the mainnet fixtures captured on
2026-08-29T09:35Z, which live in adapters/fixtures/aquarius/ and are checked into the repo, so
its arithmetic can be re-derived rather than taken on trust.
Version changelog
Every scored run is stamped with the rulebook version that produced it
(risk_scores.methodology_version, from the dex entry in METHODOLOGY_VERSIONS in
../core/src/category.ts), and it is surfaced on the API's protocol
detail and on each history point. Versions are per category and counters are independent, so this
changelog is dex's. dex v1 and lending v1 are not two editions of one rulebook and one is
not older than the other; they are two different rulebooks that each start counting at 1. The
category is stored beside the integer because the integer alone does not identify a rulebook.
| Version | Effective | Change |
|---|---|---|
| 1 | 2026-08-29 | The initial two-factor model as documented here — adminKeySafety (role posture and the upgrade reaction window) and assetControlSafety (issuer freeze and clawback) — weighted 0.55 / 0.45. depthSafety is deferred (question A, option 4) and no size floor is needed by either factor that ships. The factor set was admitted first and the weight table reviewed second; both are version 1, because no score was ever published under the factor set alone. Live since aquarius-xlm-usdc was registered on 2026-08-29; every stored dex row carries it. |
Changed under v1, without a bump: 2026-08-30, a rate limit is no longer a reading. Aquarius
was capturing a Horizon read that never reached an answer — an exhausted 429, a dropped
connection, a 5xx — as a localized cannot-assess and scoring it 0. It is now a failed run.
Recorded here rather than as a version 2 because
What bumps the version says a fix that makes the
implementation match the documented rule does not bump: the affected scores were wrong under this
rulebook, not produced by a different one. No formula, tier or weight moved. Full reasoning in
A rate limit is not a reading.
One row, and it describes the rulebook every stored dex score was produced under. It was published
before any of them existed, which is deliberate: Gate 7 requires a category to arrive at version 1
rather than acquire a version once it starts producing numbers, so the counter exists from the moment
the rulebook is published. The first stored dex row carried version 1 — settling the weight
table did not bump it, because there was no prior published number for a weight to make incomparable:
a rulebook that could not compute a score cannot have produced one that a weight made non-comparable.
From that first stored row onward the ordinary rule in
What bumps the version applies without exception,
and the next change to either weight is a version 2.
History is not backfilled across a bump, and cannot be, here as everywhere: risk_scores
stores only outputs, never the raw on-chain inputs a run was computed from.
Factor weights
Dex's weights, and dex's only. This table is the published face of CATEGORY_FACTORS.dex in
../core/src/weights.ts, which is where the adapter reads them from — no
adapter contains a weight of its own, and core/src/scoring.test.ts parses this table and fails if
the two disagree in either direction.
| Factor | Weight |
|---|---|
adminKeySafety | 0.55 |
assetControlSafety | 0.45 |
| Total | 1.00 |
Worked example — the Aquarius XLM/AQUA constant-product pool
CCSY43EHJAHT3NQDYKAMJXRFBEEH7OXDL3J3VNGO33UUSEXWNN27GBIZ, from the mainnet fixture captured
2026-08-29T09:35:43Z (adapters/fixtures/aquarius/constant-product-mainnet.ts), dex methodology
v1:
10×0.55 + 70×0.45 = 37.0 → 37
Both terms, derived from that fixture's own fields so the arithmetic is checkable rather than asserted:
adminKeySafety = 10. Role posture is the minimum over the seven roles
(§1):
| Role | Read from the fixture | Base | Activity penalty | Score |
|---|---|---|---|---|
Admin | signerCount 3, high_threshold 2 | 90 | −30 (96 ops) | 60 |
EmergencyAdmin | signerCount 1, high_threshold 0 | 40 | −0 (0 ops) | 40 |
EmergencyPauseAdmin | signerCount 1, high_threshold 0 | 40 | −0 (0 ops) | 40 |
PauseAdmin | signerCount 1, high_threshold 0 | 40 | −0 (0 ops) | 40 |
OperationsAdmin | signerCount 1, high_threshold 0 | 40 | −0 (0 ops) | 40 |
RewardsAdmin | signerCount 1, high_threshold 0 | 40 | −30 (200 ops) | 10 |
SystemFeeAdmin | signerCount 1, high_threshold 0 | 40 | −30 (200 ops) | 10 |
rolePosture = min(…) = 10. The fixture's upgrade.deadline is 0n on both the pool and the
router, so upgradeCeiling = 100 and the ceiling does not bind: min(10, 100) = 10.
assetControlSafety = 70. The minimum over the pool's two reserve tokens
(§2):
| Reserve token | Read from the fixture | Score |
|---|---|---|
XLM CAS3J7GY… | SAC, issuer.status = noIssuer — no issuer account exists | 100 |
AQUA CAUIKL3I… | SAC, issuer GBNZILST…, all four flags false, not authImmutable | 70 |
min(100, 70) = 70.
The other three fixtures, for the spread — same weights, same formulas, computed the same way,
from the same 2026-08-29 capture. All four pools read the same admin posture, so
assetControlSafety is the only term that moves:
| Fixture | Pool | adminKeySafety | assetControlSafety | Score |
|---|---|---|---|---|
constant-product-mainnet | XLM/AQUA | 10 | 70 | 37 |
concentrated-mainnet | XLM/AQUA | 10 | 70 | 37 |
stable-mainnet | USDC/USDx/yUSDC | 10 | 40 (Circle USDC is auth_revocable) | 24 |
wasm-token-mainnet | USDC (SAC) / USDC (wasm) | 10 | 40 (the wasm token is excluded, route (a)) | 24 |
The weights are an unvalidated judgment call, not an external fact.
adminKeySafetycarries more because its subject is the pool's own code and role set: a compromise there reaches every token the pool holds, the fee it charges, whether it trades at all, and the code path an LP's withdrawal runs through.assetControlSafetyreaches only the balances of one issuer's own asset and cannot touch the pool's code — a pool holding one flagged token and one clean one has a fraction of its value exposed, and the AMM's withdraw path still works.It is close behind rather than a minor term because it is the one failure Aquarius cannot mitigate and the LP gets no warning for: a clawback is a single issuer transaction with no window at all, while a code change is announced by
UpgradeDeadlinebefore it lands. Those two arguments nearly cancel, which is why the gap is 0.10 rather than lending's spread — the ordering is the claim, and the gap is deliberately small.Direction of error: if the split is wrong it is wrong by under-weighting issuer control. A reader who holds that unmitigable-and-unannounced should dominate would push toward 0.45 / 0.55, and there is no external framework that anchors against them. Full label and the rest of the list in Unvalidated judgment calls.
They were not chosen to make the scores spread out, and must not be. All 340 pools shared one admin posture in the 2026-08-27 census, and the four pools captured on 2026-08-29 still did — so
adminKeySafetydiscriminates between Aquarius markets not at all today andassetControlSafetycarries the whole variance, as the four-fixture table above shows. Down-weighting the constant factor to widen the registry's range would be calibrating a category rulebook against one protocol's current data, which is the failure this document refuses everywhere else. A weight states which failure matters more, not which reading varies more.
Both steps of the two-step admission are now done. ../TAXONOMY.md admits a
category in two reviews, and this one passed through both:
| Step 1 — the factor set | Which failures are scored, what each factor's on-chain anchor is, what was rejected. Reviewable on its own, and reviewed. |
| Step 2 — the weight table | What each factor's share is and what each formula computes, argued as judgment calls and labelled as such. Done. |
Between the two, CATEGORY_FACTORS.dex declared status: 'pendingWeights' and its factor entries
carried no weight property at all — not a zero, not a placeholder — so reading a weight off one
was a compile error rather than an undefined that would become NaN inside scoreFactors and
publish a confident-looking 0. It now declares status: 'published', and the table above is
pinned against it by core/src/scoring.test.ts in both directions: a weight edited here alone, or
there alone, fails.
The two dex factors
For each factor: the exact raw on-chain data that feeds it, what it detects, and why it is here. Every anchor names a contract and a method or storage field, per Gate 8.
Two, not five, and the missing three are not an omission. Three of lending's five have no referent in a spot AMM at all:
| Lending factor | Status for a spot AMM | Why |
|---|---|---|
utilizationSafety | No referent | It grades distance from a protocol-declared borrow cap. An AMM has no borrow ledger and no cap; nothing resembling one appears in any of the three pool wasms. |
liquiditySafety | No referent | (supplied − borrowed) / supplied is identically 1 for every AMM pool. The formula would publish 100 for all 340 pools, from no information. |
oracleSafety | No referent | Aquarius reads no price feed anywhere — established exhaustively below, not by failing to find one. Price comes from reserves, or from tick state. |
collateralSafety | Misleading if reused | HHI over reserves. A two-token pool's reserve split is its price, so the HHI is dominated by the two assets' unit prices rather than by any concentration risk — it would grade a pool worse for pairing assets of unequal price. |
adminKeySafety | Applies, same key | "Who can change the rules" is genuinely the same question. Kept under the same name, computed from different data — see below. |
Aquarius reads no oracle, and this was established rather than assumed. The exported function
list of the router, all three pool wasms, the plane, the liquidity calculator, the config storage
and the reward-boost feed were each read out of the deployed wasm. None contains a price read,
and no price-feed address appears in any instance storage. The one contract whose name suggests
otherwise — RewardBoostFeed CBKCROE56TU2FTT3C5CVN676PYVLTOQUQDHHH57GLWDY5VOKSCZPGOFN — exports
exactly total_supply() and set_total_supply(operations_admin, …) and holds
TotalSupply = 642372689091226311. It is a locked-AQUA supply feed for reward boosting; it is not a
price oracle. So oracleSafety is not "ungradable" in Gate 2's sense — it has nothing to be
ungradable about, which is why it is absent rather than disclosed.
Gate 0, re-argued for two factors
The submission proposed three factors and this rulebook ships two: adminKeySafety and
assetControlSafety. depthSafety is deferred — see
question A. Question A's
option 4 said plainly that "a two-factor category is thin enough that Gate 0 should be re-argued
before accepting it," so it is re-argued here rather than assumed to carry over.
Gate 0 asks whether the category's factors answer a question no existing category's factors already answer. Each of the two does, on its own:
adminKeySafety— the two-step upgrade reaction window. Aquarius's upgrades arecommit_upgrade→ wait →apply_upgrade, and while a deadline is pendingUpgradeDeadline − nowis exactly how long an LP has to withdraw before the code under their money changes. Lending'sadminKeySafetycannot see this: it reads one admin account's signer set and threshold, which says who could act and says nothing about how much warning anyone gets. The factor shares lending's key because it is the same question — but a lending market's answer to it is computed from a different quantity, and the reaction window is a failure mode lending's version has no access to.assetControlSafety— issuer-level freeze and clawback. A pool can be created permissionlessly against any token; if that token is a SAC whose issuer hasauth_revocableorauth_clawback_enabledset, the issuer can freeze or seize the pool's balance and an LP's exit stops depending on the AMM's code at all. No lending factor detects this, and it is not a rephrasing of one:collateralSafetymeasures concentration among assets, not whether a third party outside the protocol can take them.
Two is enough to clear the gate, because Gate 0 is a per-factor test and not a headcount. It
fails a factor that duplicates an existing one; it does not set a minimum. Both survivors detect
failures lending's five cannot see, from readings with real live variance — Circle's USDC issuer has
auth_revocable: true, a second asset also called USDC sits in the same registry, the owner key is
a 2-of-3 multisig while six other roles are lone keys, and none of that is visible from any lending
factor.
What this does NOT claim, stated because the omission is the part that could mislead. Leaving
depthSafety out is not a judgment that execution cost is unimportant to DEX risk — it is
close to the most important thing about an AMM, and the category's own Gate 0 sentence asks whether
a trader can get out at the size they are actually trading. This rulebook currently answers the
second half of that sentence (can an LP's capital leave on terms the pool's own code decides) and
not the first. The reason is narrow and it is about anchoring, not importance: Aquarius
publishes no unit of value, so any trade size we simulated at would be a number Stenion chose
rather than one the chain states, and Gate 1 exists to keep exactly that out of a published score.
The consequence is that a dex score is a governance-and-asset-control score, not a liquidity
score, and it must not be read as evidence that a pool is deep. Two things follow, and both are
obligations rather than caveats:
- Depth is still published, as a route-(a)
value: nulldisclosure carrying theestimate_swapreadings — so a reader gets the measurement without it being graded. A deferred factor is not a hidden one. - The scores will cluster, and that is a property of the data rather than a defect. All seven
roles read identically across the router and all 340 pools on 2026-08-27, so
adminKeySafetydoes not currently discriminate between Aquarius pools at all;assetControlSafetyis the only factor that varies, on the tokens a pool holds. Two pools with the same tokens will publish the same number. Anyone registering more than a handful of Aquarius markets should read that as the registry reporting the truth — these pools really do share one admin posture — and not as a ranking.
Deferred, and not declared: depthSafety
depthSafety is not one of this category's factors. It was proposed as the third and is
deferred by the resolution of
question A. It is not
declared in CATEGORY_FACTORS.dex, is not weighted, and is not computed: a declared factor with no
formula is a promise the rulebook does not keep, and weights.test.ts asserts the key is absent so
it cannot be added back without also publishing the rules for it.
What follows is kept, in full, because the deferral is about one missing input and not about the measurement being wrong. Everything here holds the day a unit of value exists; discarding it would mean re-deriving it later from a repository that no longer exists.
Anchors, when it lands: estimate_swap(in_idx: u32, out_idx: u32, in_amount: u128) -> u128 on
the pool contract, with get_reserves() and get_fee_fraction() on the same contract.
The failure it would detect: a trader cannot get out at the size they are actually trading. Nothing in lending's five measures execution cost, because a lending market has no execution — a withdrawal is at par or it is refused. That failure is real and this category does not currently measure it — see Gate 0, re-argued for what is and is not being claimed by leaving it out.
estimate_swap is a pure simulation call that runs the pool's own curve — constant product for
standard, Curve-style stableswap for stable, the tick walk for concentrated. So the slippage
figure is computed by the contract being scored rather than modelled by us, which is the strongest
form Gate 8 admits: there is no Stenion-side curve implementation to drift from the deployed one.
Both trade directions are simulated and the worse direction is the one that counts, on the same
convention every lending factor uses for reserves — the binding constraint is what a reader needs.
The fee is inside this number already, via get_fee_fraction(), because a round trip pays it. That
is the only place fee belongs — see Fee tier as a factor below.
Live readings on the XLM/USDC pair, 2026-08-27, showing the discrimination the factor exists to produce — two pools on the same pair, one materially deeper, both figures produced by their own contracts:
| Pool | Type, fee | 1,000 XLM in | 1,000,000 XLM in | Impact |
|---|---|---|---|---|
CA6PUJLBYK… | standard, 10 bps | 0.186277 USDC/XLM | 0.168205 | −9.70% |
CBBMQBNHB2… | concentrated, 10 bps | 0.186265 | 0.175584 | −5.74% |
Cannot-assess would resolve to 0, the unsafe end. A pool whose estimate_swap reverts, or whose
reserve is zero, would score 0 — never 100, and never a skipped factor, matching
collateralSafety's existing treatment of an empty set.
What is NOT settled, and is why this is deferred: the trade size it simulates. Everything above describes how the cost is measured, and all of it holds. At what size has no answer in Aquarius's own terms — a pool-relative size makes the factor degenerate across 272 of 340 pools, and an absolute size needs a unit of value Aquarius does not have. See question A. No adapter may implement this factor until that changes.
1. adminKeySafety — seven roles, and how long you get to react to a code change (weight 0.55)
Anchors: get_privileged_addrs() -> Map on the router and on every pool; the UpgradeDeadline
and FutureWASM entries in contract instance storage; Horizon /accounts/{G…} for each role's
signer count, thresholds and recent activity.
This is the same factor key lending uses, and that is a decision, not an oversight. It is recorded here so it is not re-litigated, and it was an open question when the category was admitted:
Gate 0 rejects a factor that is "an existing factor rephrased, rescaled, or renamed for a new audience." "Who can change the rules under my money" is not a rephrasing of lending's question — it is the same question, and the honest way to say so is to use the same name. The data underneath is entirely different: lending reads one admin account's signer set and threshold, while this reads seven named roles plus a two-step upgrade deadline. A separate key —
roleControlSafety, say — would have published two names for one failure and started the taxonomy fragmenting into synonyms, which is the drift the sharedRiskFactorTypevocabulary exists to prevent. So: one key, two computations, declared per category inCATEGORY_FACTORS.A shared key is not a comparability claim. A
dexadminKeySafetyof 60 and alendingadminKeySafetyof 60 were produced by different rules from different data and mean different things — exactly as the two categories' overall scores do. See Comparability.
The failure it detects that lending's version cannot: the upgrade reaction window. Aquarius
upgrades are two-step. commit_upgrade(admin, new_wasm_hash, …) writes UpgradeDeadline;
apply_upgrade(admin) refuses until that deadline passes; revert_upgrade(admin) cancels. When a
deadline is pending, UpgradeDeadline − now is exactly how long an LP has to withdraw before
the code under their money changes. It is anchored with no Stenion constant in it at all — the
chain states the deadline and the chain states the time.
Read on 2026-08-27: UpgradeDeadline = 0 and FutureWASM == running wasm on the router and all
340 pools. No code change is scheduled anywhere.
The role structure, read on 2026-08-27 — identical across the router and all 340 pools:
| Role | Account | Signers | high_threshold | ops in 30d |
|---|---|---|---|---|
Admin (owner / upgrade) | GAV5FBMKD2ZF4X2MGWDNQYUP7KFL7MRM6HZBY7HKQLB4BRHSCCX5J6VS | 3 | 2 | 96 |
EmergencyAdmin | GCGZ6E5RBUKLNB4VZ5RC65C4QMBSBJ3COVRRJCWAMCXJC36LB7YYWEKM | 1 | 0 | 0 |
EmergencyPauseAdmin | GA6MVTGQDCJPP27IAMG6PSDTWOJYTD3NUTLR2W54ADBCBY7OID5YUDSI | 1 | 0 | 0 |
PauseAdmin | GA6MA665XVKHTQUZVSMUKUPGT7OREJNCLAZ5ZEH5CXPKYTWFJKZ3YSEK | 1 | 0 | 0 |
OperationsAdmin | GBVQPX2LQ55HLRMLIWBEYVVQL3SZ5RFPRKYLRSLZU4XRIWAXW2KQIMMD | 1 | 0 | 0 |
RewardsAdmin | GCXYKA3BM574WC6TWESEDUGUJTNQ5SVCFMHWLQ634H5FTE7FYPV3JH3X | 1 | 0 | 200 |
SystemFeeAdmin | GB57YDVGLL2BAVOXHPXYCZR77J4MLPLMGJKFMTUKMHFI2AEGS4SGGW7N | 1 | 0 | 200 |
The owner key is a 2-of-3 multisig; the other six are lone keys. That is a real, checkable posture and it is not the one Aquarius's own auditor recommended — Certora's Appendix A asks for the Pause Admin to be "a multisig or DAO" and for the owner key to be kept offline. Both halves are readable, and both disagree with the recommendation. Publishing that disagreement is the factor doing its job.
Formula — a per-account tier, minimised over the roles, then capped by the upgrade state.
Two components, each computed from the readings above, combined by taking the worse of the two.
Component 1 — role posture, from get_privileged_addrs() and Horizon /accounts/{G…}:
base = 90 signerCount > 1 AND high_threshold > 1 # N-of-M multisig
40 a classic account that is not a multisig # single master key
0 a `C…` contract address, or the lookup failed # cannot assess -> unsafe end
activityPenalty = min(30, recentOps × 3) # recentOps = operations in the last 30 days
accountScore = clamp(base − activityPenalty, 0, 100)
roleScore = min(accountScore) over the addresses that role holds
rolePosture = min(roleScore) over the seven roles
Component 2 — the upgrade reaction window, from UpgradeDeadline and the fetch timestamp:
upgradeCeiling = 100 UpgradeDeadline == 0 # no code change is scheduled
40 UpgradeDeadline > fetchedAt # scheduled; the window is still open
0 UpgradeDeadline <= fetchedAt, non-zero
# matured: applicable at the next
# ledger, with no warning left
adminKeySafety = min(rolePosture, upgradeCeiling)
Why these numbers:
- The per-account tier is lending's, adopted rather than re-argued —
90for an N-of-M multisig,40for a single master key,−3per operation capped at−30, exactly as lending §3 publishes them. The reading underneath is identical on both sides: Stellar's account threshold model and a 30-day operation count, from the same two Horizon calls. Inventing a second set of integers for the same reading would be two rulebooks for one question — which is the drift the sharedadminKeySafetykey exists to prevent, applied to the numbers rather than to the name. So the split is anchored exactly as it is there (a 1-of-1 key is a single point of unilateral compromise;signerCount > 1 AND high_threshold > 1provably is not, per Stellar's own threshold model), and the exact integers remain a labelled judgment call exactly as they are there. - The contract-address branch is
0here and60in lending, and that divergence is the one deliberate departure. Argued in Cannot-assess below: a contract admin is a known, named structure lending chose not to grade, while an Aquarius role that is not a classic account is a structure we did not expect and cannot describe. - The
40inupgradeCeilingis not a new constant — it is the single-master-key tier, reused. While a code change is scheduled, the LP's only remaining protection is a countdown whose length cannot be read at all, so the market is graded no better than one whose rules a lone key can change. The rule introduces no integer this document did not already publish. - The
0once the deadline has matured is a reading, not a choice. What this component grades isUpgradeDeadline − now, and once that is non-positive the measured warning is zero. The chain states the deadline and the chain states the time; there is no Stenion constant in that branch. FutureWASMis read but does not enter the formula. Every contract read on 2026-08-29 carried aFutureWASMequal to its own running hash, so presence is the quiescent state rather than the signal —UpgradeDeadlineis what says a change is scheduled.FutureWASMand whether it differs from the running hash are published in the factor'sdetailstring, because a reader wants to know which code is staged, and they are not graded because the deadline already carries the fact.
Two of the choices above are unvalidated judgment calls — combining the seven roles by min,
and the pending-upgrade ceiling. Both are labelled, with their direction of error, in
Unvalidated judgment calls.
Two limits, stated rather than papered over:
- The timelock duration is not readable, and that is why
upgradeCeilingis a state rather than a curve.ADMIN_ACTIONS_DELAYis a compile-time constant with no getter, and the source repository is gone. The chain states the deadline and the chain states the time, so the factor can say whether a window is open, say when one has matured, and publish the seconds remaining — what it cannot do is say what fraction of the window is left, because it never learns the whole. Any grading of the remaining window's length would need a threshold in seconds that Aquarius does not state, which is an invented number and ground rule 4 forbids it. So the duration itself takes Gate 2 route (a): avalue: nulldisclosure component saying it is not readable from the contract, published beside the factor rather than folded into it. - The Emergency Admin can bypass the delay. In Aquarius's own words, in their response to Certora's H-01: "In the case of system vulnerability fixes, delay may be bypassed by the Emergency Admin role." So the reaction window is conditional on one key choosing not to skip it — and that key is one of the six single-signer accounts above. Also route (a): disclosed beside the window, never silently folded into it, because "a window exists" and "the window is unconditional" are different claims and only the first is true.
Why route (a) and not the other two, for both, as Gate 2 requires be argued rather than
asserted. Route (b) — a precondition that leaves the protocol unscored — would refuse to score
every Aquarius market over two facts that are the same for all 340 of them and that neither
changes nor discriminates: nothing would ever be published, and a reader would learn less, not
more. Route (c) — live ungraded state — is for readings that change between cycles, and neither
does: ADMIN_ACTIONS_DELAY is a compile-time constant in code we cannot read, and the Emergency
Admin's bypass is a property of the deployed contract's authorisation logic. Both are fixed facts
about an unreadable quantity, which is exactly the shape route (a) exists for.
2. assetControlSafety — can a third party freeze or seize what the pool holds (weight 0.45)
Anchors: each reserve token contract's instance executable (a Stellar Asset Contract's is
contractExecutableStellarAsset) and its METADATA instance entry, whose name is CODE:ISSUER
for a classic asset and the bare string native for XLM, then Horizon /accounts/{issuer} for that
issuer's flags.
Corrected from the submission, which named an
AssetInfostorage key. There is no such key; The issuer was found inMETADATAby probing the deployed token contracts. The Gate 8 claim is unchanged — the issuer is read from the token contract's own instance storage, not from an aggregator — only the key name was wrong. Thenativecounter-example gets sharper as a result: XLM'sMETADATA.nameis the bare stringnative, which is why a SAC is detected from itsexecutableand never from the shape of its name.
The failure it detects, and it has no analogue in lending's five: an Aquarius pool can be
created permissionlessly against any token. If that token is a Stellar Asset Contract whose issuer
has auth_revocable or auth_clawback_enabled set, the issuer can freeze or seize the pool's own
balance — and an LP's exit stops depending on Aquarius's code at all. No amount of depth and no
admin posture protects against it. Aquarius's auditor named this too: Certora M-02, "Lack of scam
protection for AMM Users."
Readable, with real variance. Of 205 distinct tokens across the 340 pools, 196 are Stellar
Asset Contracts (executable = contractExecutableStellarAsset, issuer parsed out of METADATA)
and 9 are wasm contracts. Sampled issuer flags, 2026-08-27:
USDC:GA5ZSEJY…(Circle,home_domain = circle.com) —auth_revocable: true. Same forEURC:GDHU6WRG….LSP:GAB7STHV…—auth_immutable: true, the strongest reading available: the flags can never change.- AQUA, yXLM, USDx, EURx, BLND, SSLX, WHLAQUA, XRF — all four flags false.
- And a second asset also called
USDC, issuerGCBYVQH3…,home_domain = mirrasets.com, sitting in the same registry as Circle's. That is Certora M-02 in the live data, not in theory.
Formula — a per-token tier on the issuer's flags, minimised over the pool's reserves:
Per reserve token:
100 a SAC with no issuer account at all — native XLM
100 auth_immutable set, with auth_revocable AND auth_clawback_enabled both clear
70 auth_revocable and auth_clawback_enabled both clear, auth_immutable not set
40 auth_revocable set, auth_clawback_enabled clear
0 auth_clawback_enabled set
0 the token IS a SAC and the Horizon read for its issuer failed
- not a SAC: excluded from the computation, route-(a) disclosure (see below)
assetControlSafety = min(tokenScore) over the graded tokens
0 when no token is gradable
The order the tiers are tested in is load-bearing. Clawback is tested before revocable, and both
before immutable, because auth_immutable freezes whatever the flags currently are — it is a
credit only when what it freezes is clean. An issuer that is both immutable and revocable has made
its freeze power permanent and scores 40, not 100. Testing immutability first would invert that.
auth_required moves no number, deliberately. It gates who may acquire the asset; it does not
let an issuer touch a balance that already exists. The power to freeze one is auth_revocable and
the power to take it is auth_clawback_enabled. auth_required is read and published in the
factor's detail string, because it describes the asset a reader is looking at, and it is not graded
because it does not answer this factor's question.
Why these numbers. The ordering is anchored to what Stellar's account flags let an issuer do,
which is a protocol-level fact and not a Stenion preference: seizure (auth_clawback_enabled)
strictly dominates freezing (auth_revocable), which strictly dominates an issuer holding neither,
which is in turn a weaker statement than an issuer that can never acquire either — and weakest of
all against an asset with no issuer account for anyone to act from. The spacing — the 40 and
the 70 — is an unvalidated judgment call, labelled in
Unvalidated judgment calls.
Why 70 and not 100 for an issuer whose flags are clean. Because auth_revocable is one
SET_OPTIONS away for any issuer that has not set auth_immutable, so "clean today" is a strictly
weaker statement than "cannot become dirty". That distinction is the whole reason the second asset
also called USDC (GCBYVQH3…, home_domain = mirrasets.com) is worth publishing about: its flags
read exactly like a well-run issuer's, and in both cases the reading is one transaction from
changing. Scoring it 100 would publish the reading as a guarantee.
The 9 wasm tokens have no issuer-flag equivalent (SolvBTC, xSolvBTC, BnUSD, XAUM, and wasm
USDC/USDT variants). They take Gate 2 route (a): a value: null disclosure component naming the
token and saying the read does not apply. Not a silent pass, and not a silent 0.
Why not the other two routes, as Gate 2 requires be argued rather than asserted: route (b) — leaving the protocol unscored — would drop every pool containing any of nine tokens over a signal that is absent rather than broken, which is far stronger than the finding warrants. Route (c) — live ungraded state — is for state that changes, and "this token is not a SAC" is a fixed property of the contract's executable, not a reading that can flip between cycles.
Cannot-assess, stated per factor — and what is a failed run instead
Gate 2 requires every "cannot assess" branch to resolve to the unsafe end of the scale. That is one sentence in the checklist, and it needs an answer per factor rather than one example standing in for the set — so both factors state their own below, and neither inherits the other's by implication.
The rule, stated once and applied to both: a partial or unreadable input scores 0, never a skip
and never a pass. Where a read leaves any part of a factor's input unavailable — a call that
reverts, a response shorter than expected, a lookup that fails for one subject — the factor resolves
to 0, the unsafe end. It is never omitted from the factor map, never defaulted to a neutral
middle, and never allowed to score well because there was less to check. This is ground rule 4, and
lending's 2026-08-16 correction is the precedent: two factors there returned 100 from an empty
filtered set, which published "maximally safe" from no data, and both were changed to return 0.
One boundary, and it is the difference between publishing a 0 and publishing nothing. A whole-endpoint outage — Soroban RPC unreachable, Horizon down, nothing decodes — makes the adapter throw, and the indexer records a failed run and publishes no score for that cycle (CLAUDE.md's error-handling rule; error handling lives in the indexer, never per adapter). That is not a cannot-assess branch and must not be graded 0, or a blip in our own network path would publish "dangerous admin control" across every pool at once. Everything localized — one call, one role, one issuer — is a cannot-assess branch and takes the 0.
"Localized" is not a count of failed reads; it is a claim about what the failure is a statement
about. The two words that separate the branches are the subject answered. A read whose subject
answered — an issuer account Horizon says is not there, a contract that reverted — produced a fact
about the pool, and a fact about the pool is what a factor is for. A read that never reached an
answer — the endpoint refused us for rate, the connection dropped, Horizon returned a 5xx about
itself — produced a fact about Stenion's read path on that cycle, and no such fact may move a
protocol's number. Both used to be called "the lookup failed"; they are opposite claims, and the one
that is not about the protocol fails the run. The argument, the options weighed against it, and what
it does and does not change are in
A rate limit is not a reading below.
| Factor | Cannot-assess branch | Resolves to |
|---|---|---|
adminKeySafety | get_privileged_addrs() reverts; returns fewer than the seven roles; a role is ungradable; the contract carries no UpgradeDeadline entry at all | 0 |
assetControlSafety | Horizon answers about a SAC issuer and the answer is not a gradable account; no reserve token is readable at all | 0 |
| (either factor) | The endpoint never answered: a rate limit we could not wait out, a dropped connection, a 5xx | Not a branch — a failed run, and no score published for the cycle |
| (disclosures) | Timelock duration; Emergency Admin bypass; the 9 non-SAC wasm tokens | Not a branch — route-(a) value: null, moves the number neither way |
adminKeySafety. Three ways the input can be incomplete, all resolving to 0:
get_privileged_addrs()reverts on the router or on the pool. The whole role structure is the factor's input, so a revert leaves nothing to grade. 0, not a skipped factor — "we could not read who controls this pool" is a statement about the pool, and the unsafe end is the only honest place for it.- It returns fewer than the seven expected roles. Seven were present across the router and all
340 pools on 2026-08-27 (
Admin,EmergencyAdmin,EmergencyPauseAdmin,PauseAdmin,OperationsAdmin,RewardsAdmin,SystemFeeAdmin). A short map means either an unexpected contract version or a role we cannot see, and grading the roles that did come back would publish a posture assessment of an admin set we know is incomplete. 0. - A role reads but cannot be graded — its address is not a classic
G…account, so there is no signer set, threshold or activity history behind it. 0. - The contract carries no
UpgradeDeadlineentry at all — the raw shape records this asdeadline: null, which is a different statement from0nand must not be collapsed into it: one says the contract answered "nothing pending", the other says this contract does not keep the field the upgrade component reads. Both the router and pools of all three types carried it on 2026-08-29, so its absence would mean an unexpected contract version, and half this factor's input would be missing. 0, on the same reasoning as a short role map.
That third case argues with an existing precedent, and the resolution is recorded rather than assumed. Lending's
adminKeySafetysends its equivalent to a clearly-flagged neutral 60, on the reasoning that a contract-held admin is a different structure rather than evidence of danger.dexdoes not inherit that baseline: lending's 60 is anchored to a contract admin being a known, named structure whose properties we chose not to grade, whereas an Aquarius role that is not a classic account is a structure we did not expect and cannot describe. A neutral score for an unexpected reading is an invented number, which ground rule 4 forbids. All seven roles were classic accounts on 2026-08-27, so this branch is unexercised today.
assetControlSafety. Two cases, and the distinction between them is the whole point:
- Horizon answers about a SAC issuer's
/accounts/{issuer}and the answer is not a gradable account — a404, most plainly. The token is a SAC, so its issuer flags are exactly the thing this factor grades, and Horizon has told us there is nothing there to grade. 0 for that token, which then binds the factor on the usual worst-reserve convention. This is not the same as the 9 non-SAC tokens and must not be routed like them: there, the read genuinely does not apply; here, it applies and failed. Treating a failed read as "does not apply" would silently upgrade an unknown into an exemption, which is the single most dangerous confusion available in this factor. It is equally not the same as a read that never reached an answer — see A rate limit is not a reading. - No reserve token is readable at all — every reserve is a wasm contract, so every token takes
the route-(a) disclosure and nothing is left to grade. A minimum over an empty set is 0, not 100. Same shape as lending's
2026-08-16correction, written down before the adapter exists rather than after it publishes a 100.
The 9 non-SAC wasm tokens remain what they were: a route-(a) value: null disclosure naming the
token and saying the read does not apply. A disclosure is not a zero and not a pass — the token is
excluded from the computation, not graded badly by it.
A rate limit is not a reading
The question. Aquarius is the first adapter that scores through a failed read instead of throwing on it, and the 429-backoff work made the consequence concrete: when Horizon refuses us for rate and the attempt's retry budget is spent, the refusal arrived at scoring as a 0 with a reason string, indistinguishable at the factor level from "this issuer's flags could not be read." Those are different claims. One is about the pool. The other is about Stenion's infrastructure at the moment of the read, and it was moving a published number.
The decision: a rate-limit exhaustion, and every other failure that never reached an answer, is a failed run — not a scored 0. Nothing about the two formulas, the tier values, the weights or the worst-reserve convention changes. What changes is which failures are allowed to reach them.
The rule, in the form the code applies it. A read is a reading only when its subject answered, and each transport is asked that question in its own vocabulary:
| Transport | An answer about the subject — captured, scores | Everything else — fails the run |
|---|---|---|
| Horizon | an HTTP status that is about the account: 404, and any other 4xx | a 429, any 5xx, a rejected fetch, a body that did not parse |
| Soroban RPC | a simulation error — the contract itself refusing | a rate limit, a dropped connection, an SDK decode failure |
The classification lives in ../adapters/read-failure.ts, not in
each call site, and it is tested there and in adapters/aquarius/fetch.test.ts against a loopback
Horizon answering 404, 503, 429 and nothing at all.
Why, in three facts rather than a preference.
- A 429 carries no information about the protocol. Every one of Aquarius's 340 pools is equally rate-limitable, on a schedule set by our own request train against a free shared endpoint. A number that moves with it is reporting Stenion's cycle, and the registry has no column for that.
- It is not localized, which is the word the boundary above turns on. The retry budget is
per attempt and shared by every call the attempt makes (
../core/src/rate-limit.ts): once it is spent, the next refused read gets no retries at all. So one exhaustion silences the reads after it,adminKeySafetyandassetControlSafetyboth resolve to 0, and the pool publishes an overall 0 — "dangerous admin control", from an endpoint that was busy. That is precisely the outcome the whole-endpoint boundary exists to forbid, arriving one pool at a time instead of all at once. - The honest publication already exists and costs nothing to reach. A throw runs the
indexer's own target retry with a fresh budget, up to
STENION_RETRY_ATTEMPTStimes; only if every attempt is refused does the cycle record a failed run, flaggedrateLimitedand carrying the 429s inrisk_scores.error. The previous score stands with its staleness advancing, which says the true thing: we did not learn anything about this protocol this cycle.
The options that were weighed, and why each was not taken.
| Option | Why not |
|---|---|
| Keep the scored 0. Gate 2 resolves uncertainty to the unsafe end and does not ask why. | Gate 2 governs a cannot-assess branch, and a cannot-assess branch is a reading. The rule it appeals to is ground rule 4 — never publish a number no data supports — and a 0 sourced from our own request rate is that number, not an application of it. |
| Carry the prior cycle's sub-reading forward and retry next cycle. | risk_scores stores outputs, never the raw inputs a run was computed from, so there is nothing to carry: this needs a new persistence layer for per-sub-value state, and a score assembled from two cycles' readings is a fourth publication route nobody has argued for. Flagged as a separate design question, not forced into this pass. |
Keep the 0 and disclose it in detail, plus a dashboard treatment. | It fixes the ambiguity for a reader who opens the factor and not for the ranking, which is where the damage is: the pool sits last in the dex block with a footnote. Disclosure is the right response to something we chose not to grade; this is something we never measured. |
| Throw on a 429 only, leaving the other never-answered failures scoring 0. | A dropped connection and a Horizon 5xx are the same claim about the same thing and were scoring 0 too — "Horizon down" was publishing a risk finding, against the boundary this document already wrote. Fixing the named case and leaving its siblings would have made the rule a special case instead of a rule. |
What this deliberately does NOT do. Aquarius still scores through every failure that is a
statement about the pool: a reverting get_privileged_addrs(), a short role map, an ungradable
role, an issuer account Horizon says is not there. The score-through-failure design is intact — it
was never a design to score through our own failures. Soroban RPC's 429 handling is unchanged;
what changed is that Aquarius no longer swallows the RateLimitExhaustedError that handling
produces. Lending needed none of this and is untouched: Blend and Kinetic throw on every failed
read, so no fact about our read path could reach one of their numbers in the first place.
No version bump, deliberately. Per
What bumps the version, a fix that makes the
implementation match the rule already documented here does not bump — the affected stored scores
were wrong under this rulebook, not scored under a different one. No formula, threshold, weight or
tier moves, and a dex score computed from readings is byte-identical before and after. What
changes is that some cycles now publish no number where they previously published a wrong one, and
risk_scores.methodology_version was never the field that distinguished those.
Every branch above is implemented and tested, and the adapter may not reinterpret it.
The rules were written before the code, which is the order ../TAXONOMY.md
requires; adapters/aquarius/score.test.ts now asserts each one individually, including the four
that no live pool can reach — a reverting get_privileged_addrs(), a short role map, a
contract-held role, and a pool whose every reserve is a wasm contract. All of them return 0,
each with a detail naming what could not be assessed.
And so is the line above them. adapters/read-failure.ts decides, once, which failures are
allowed to become one of those branches at all; adapters/read-failure.test.ts pins the
classification and adapters/aquarius/fetch.test.ts drives the two Horizon readers against a
loopback endpoint answering 404 (captured as the reading), 503, a refused connection and a
persistent 429 (all three a failed run, the last still recognisable downstream as a rate limit).
None of the four can be produced from live data on demand, which is why they are pinned rather than
left to be discovered on the cycle they first happen.
Size floor: none, and none pending
There is no size floor in this rulebook, and there is no unwritten one waiting to be added. Gate 4 asks for a size below which a published number stops carrying information. For the two factors that ship, no such size exists — and that is a property of what they measure, not a gap left open:
adminKeySafetyreads the same seven roles whatever the pool holds. A pool with 3 stroops in it has the sameAdmin, the same six single-signer keys, the sameUpgradeDeadlineand the same Horizon signer sets as one holding 10,000 XLM. Every input is a property of the contract and of the accounts that control it; not one of them is a quantity of anything.assetControlSafetyreads the same issuers. Whether Circle can freeze the pool's USDC does not depend on how much USDC is in it. A dust pool holding a clawback-enabled asset is exactly as exposed as a deep one.
Nothing here degrades as a market gets smaller, so there is nothing for a floor to protect.
Lending needs one for the opposite reason, spelled out in
the market-size floor: every one of its five factors falls to a
can't-assess branch on an empty market and every one of those branches is 0, so an empty market
would publish a danger-band number meaning the opposite of the truth. Neither dex factor has that
failure mode — an empty pool's roles and issuers still read, and the number they produce is still
true of it.
This is a consequence of question A's resolution, not an open item, and the distinction is
the whole point. The census that would motivate a floor is real and dated — of the 148 XLM-paired
pools on 2026-08-27, 18 held zero XLM and 59 held under 1 XLM, many of them 1–3 stroops, while
only 16 held more than 10,000 — but a floor is needed by a size-sensitive factor, and the one
that would have been size-sensitive (depthSafety) is deferred. A floor becomes necessary again
the moment depthSafety lands, which is the same moment the unit of value to express one in would
exist. Writing one before then would mean choosing a denomination for a factor that does not exist,
in a category with no unit of value — the permanent, one-protocol decision question A declined to
make.
No number is copied from lending's market-size floor or its minimum-size filter, and none may be. Those are denominated in USD against a lending pool's supplied value, which is a quantity this category cannot compute.
What this does not claim, and where an under-sized market goes. It does not claim a 3-stroop pool is worth anyone's attention — it claims that the two numbers this rulebook publishes about one are as true of it as they are of a deep pool, which is a statement about the factors and not about the market. Nothing is excluded for size, so no pool needs a coverage entry on those grounds. Which pools are registered is a separate reviewed decision about what is worth publishing, and never a statement that an unregistered pool could not be scored.
Unvalidated judgment calls
../TAXONOMY.md Gate 1 requires every threshold in a rulebook to either name the
on-chain field it anchors to or be labelled an unvalidated judgment call, in both the code
and methodology/, with its reasoning and its direction of error stated. This is the complete list
for dex v1. Every other number in either formula names a field or is itself a reading; where one
is adopted from lending rather than chosen here, that is said below rather than left to be noticed.
| Constant | What it is | Direction of error, in one line |
|---|---|---|
the two weights, 0.55 / 0.45 | each factor's share of the overall score | under-weights issuer control if wrong; the gap is deliberately small because the ordering is the claim |
combining the seven roles by min | how seven separately-read roles become one number | understates safety, never overstates it — a weak minor role binds as hard as a weak Admin |
the pending-upgrade ceiling, 40 | what a scheduled code change does to the number | too harsh for a routine announced upgrade, far too lenient for a hostile one |
the asset-control tier spacing, 40 and 70 | how far apart freeze-capable and clean-but-mutable sit | 40 leans lenient, 70 conservative — and min makes the lenient end bind, so this one reads generously |
The two weights, 0.55 / 0.45. Argued at length in Factor weights and
labelled in ../core/src/weights.ts. No external framework anchors them.
If the split is wrong it is wrong by under-weighting issuer control, and a reader who holds that
unmitigable-and-unannounced should dominate would push toward 0.45 / 0.55. The gap is 0.10 rather
than lending's spread because the two arguments nearly cancel: adminKeySafety reaches more of the
pool, assetControlSafety reaches it with no warning and no remedy.
Combining the seven roles by min. Every one of the seven can act unilaterally within its own
scope, so an attacker takes whichever key is weakest and the weakest is what the number should
report. The cost is that min treats the seven as equally consequential, which they are not — a
SystemFeeAdmin compromise is not an Admin compromise — so a pool whose only weak key is a minor
one is graded as though the code-upgrade key were weak. It therefore understates safety and never
overstates it, which is the direction ground rule 4 requires a judgment call to lean. Two
alternatives were considered:
- Weighting the roles by how much each can do. It would fix the objection exactly, and it needs
seven Stenion-chosen numbers where
minneeds none — turning a four-row type-(b) list into a ten-row one, which by TAXONOMY.md's own framing would itself be the finding. - Grading the owner key alone. Rejected outright. Aquarius's
Adminis a 2-of-3 multisig, so this would publish a comfortable number while ignoring six live single-signer keys that can pause the market, move fees and set rewards. It is the alternative that most looks like a simplification and is in fact a decision to stop reading six of the seven readings.
The pending-upgrade ceiling, 40. It introduces no integer this document did not already
publish — it is the single-master-key tier, reused on the reasoning that a scheduled code change
leaves the LP holding a countdown of unreadable length rather than a signer structure. Direction of
error, both ways: it is too harsh for a routine, correctly-announced upgrade, because it lowers
the number of a protocol for using the very two-step mechanism this factor credits it for; and it is
far too lenient if the staged code is hostile, in which case no score would be adequate. It
leans conservative, and it is unexercised today — UpgradeDeadline read 0 on the router and on
pools of all three types on 2026-08-29. The alternative was to leave the pending state ungraded and
publish it as a disclosure; that was declined because the reaction window is the failure mode this
factor's Gate 0 argument rests on being able to see, and a
factor that discloses it is a role-posture factor with a footnote.
The asset-control tier spacing, 40 and 70. Only the two interior values are chosen; the
ordering around them is anchored to what Stellar's flags let an issuer do.
Direction of error, and it is the one entry in this list that is not symmetric. 40 for
freeze-capable leans lenient: to an LP, a freeze that is never lifted is indistinguishable from
a seizure, and 40 asserts the two differ. 70 for clean-but-mutable leans conservative: it
refuses to call a mutable issuer as safe as an asset with no issuer, and so understates a
long-standing issuer that has never set a flag and never will. Those pull opposite ways, and the
lenient end is the one that binds, because the factor is a minimum over the pool's tokens: any
pool holding a freeze-capable asset is scored by the 40 and the 70 never enters the number at
all. So the net lean of this constant is toward reading generously — the opposite of the
role-combination rule two entries above, which understates safety and never overstates it, and the
same direction as the weights' possible under-weighting of issuer control. Those two
compound, and on exactly one shape of pool: one whose only weakness is a freeze-capable issuer is
graded by the lenient tier and has that tier's factor weighted at the lighter of the two. That
pool is where a published dex number is most likely to be too high, and it is the first thing to
challenge if one looks generous.
Moving the two values together makes the factor a pass/fail on clawback alone; moving them apart makes it nearly binary on any flag at all.
Adopted from lending rather than chosen here, and labelled there: the per-account tier values
(90 for an N-of-M multisig, 40 for a single master key) and the activity penalty (−3 per
operation, capped at −30), from
lending §3. They are the
same integers applied to the same Horizon reading, and that section carries their label, their
reasoning and their direction of error. They are named here so a reviewer counting this rulebook's
constants finds them, and they are not re-argued here, because a second argument for the same
integers is how two rulebooks start. Changing them on one side and not the other is a decision to
fork them and needs its own argument — it is not an edit to one file.
Four rows, where the weight-table review was expected to bring two — and the growth is stated rather than smoothed
over, because TAXONOMY.md's rule is that a long type-(b) list is itself the finding. This one is
not long, and both additions are accounted for: that prediction was written when this rulebook had
no formulas at all, only anchors, and the two new rows are precisely the two places a formula had
to say something the chain does not state — how seven separately-read roles become one number, and
what a scheduled code change does to it. Neither is a preference standing in for a measurement, and
neither introduces an integer published nowhere else. On the other side the list shrank: that
same prediction included rows for depthSafety's probe size, its impact-to-score mapping, and the size floor,
and all three are gone because the factor and the floor are gone with them.
Three of the four are labelled by the adapter's scoring code, not by this document. The weights live in
CATEGORY_FACTORS.dexand carry their label there. The other three are constants inadapters/aquarius/score.ts, which labels each one beside the branch that applies it — the per-account tiers adopted from lending by reference, the reused40in the open-upgrade-window branch, and the40/70asset-control tiers. Gate 1's "in both the code andmethodology/" is therefore discharged for all four. Recorded here so the obligation stays attached to the rulebook rather than remembered.
Live ungraded state — route (c), published beside the score, never in it
Aquarius's pause surface is real, changing, on-chain and ungradable: exactly the shape
operationalState exists for (Operational state).
Every pool exports get_is_killed_swap(), get_is_killed_deposit(), get_is_killed_claim() and
get_emergency_mode(); the router exports get_emergency_mode().
Read on 2026-08-27: router emergency mode false; two stable pools have
is_killed_deposit = true — CDKVJYMN34ZIEXSLNFYHVAFF6M6FM5E2U6OHXOTBKH2WLBULXOE53YDP (XLM/AQUA,
zero reserves) and CBKENQ33KITYE4JKPAWALHU4KGWV5AXQLJFUNRNIDGCRLRDENX6PYVDE
(AQUA/CCKCKCPH…) — everything else false, no pool in emergency mode. The field has live variance
on day one, which is what makes it worth publishing rather than a field that is always the same word.
The structural finding, and the strongest single fact about an Aquarius LP's exit risk: there is
no kill_withdraw. The exported function list of all three pool wasms contains kill_swap,
kill_deposit, kill_claim, kill_gauges_claim and their unkill_ counterparts — and no
withdraw equivalent. Withdrawals cannot be halted by any Aquarius role. That is the same class of
fact as "Blend never blocks withdrawals at any status," established the same way: it is a property
of the deployed code, not a promise.
Read the state through the getters, never through instance storage. The three pool types
disagree about key names — standard writes IsKilledClaim, concentrated writes
ClaimKilled/IsKilledSwap/EmergencyMode — and a flag that has never been toggled has no key
at all. The getters normalise all of it; raw storage would make an untouched pool
indistinguishable from one that was read wrongly.
The swapDisabled rung, and why the shared ladder gained one
dex registers its own operation vocabulary — { swap, deposit, withdraw, claim } — and its own
canonical ordering and classifier in
../core/src/operational-state.ts, beside lending's. No
lending operation was renamed, added or removed; renaming one would rewrite blocked in every
stored operational_state and every API response.
OperationalLevel stayed one shared ladder, and gained a rung rather than forking — an open
question when the category was admitted, recorded here so it is not re-litigated:
A pool with
is_killed_swap = truebut deposits and withdrawals live would have classified asactiveunder the ladder as it stood — true about exit, and wrong about the market.blocked: ['swap']carried the fact, butlevelis the field a reader scans, and a dead market reading "active" is the quiet misstatement the whole type exists to prevent. So the ladder gainedswapDisabled: cannot trade; depositing and withdrawing both work.Why not generalize
borrowingDisabledinto one sharedcoreActivityDisabledinstead, which would have been the more honest name for both:borrowingDisabledis a value stored in every historicaloperational_stateand published on the public API, so renaming it is a breaking change to stored data and to the API contract, inside a change whose entire purpose was to add a category. Two named rungs and an accuratelevelbeat one elegant rung and a migration nobody asked for. The two sit at equal severity and can never meet, sinceblockedcarries one category's vocabulary.
exitDisabled is unreachable from Aquarius — there is no kill_withdraw — and the rung stays,
because it is the top of the ladder every category is measured against and a second DEX that can
freeze withdrawals must have somewhere to say so. claim gets no rung at all: an LP whose reward
claim is killed can still withdraw every unit of principal. It stays in blocked because it is
true, on the same reasoning lending gives for repay and liquidate.
What was rejected, and why it must not be re-proposed
Gate 3. Every candidate that was considered and dropped, by name, with its reason — and with the measurement where one was taken.
liquidity_calculator.get_liquidity() as a size or depth measure
Aquarius ships its own cross-pool liquidity metric (CCUKQWLM…, get_liquidity(pools) -> Vec),
used to allocate AQUA rewards. It is the obvious thing to reach for, it is on-chain, and it clears
Gate 8 — and it is not comparable in any economic sense. The evidence is a single pool: on
2026-08-27, CBMSBM6EABGBNZ47WZTLX6WJOB3ETO2GNPQYC2FX6LJXUL7TRQFZ3IA3 — a standard XLM/DATAVAULT
pool holding 32,733 stroops, i.e. 0.0032733 XLM, roughly a third of a US cent — ranked 4th of
all 340 pools by this metric. It is dominated by raw token quantities, so a pool paired against a
huge-supply token outranks the XLM/USDC book. Fails Gate 1 (no economic anchor), and would make any
Gate 4 floor built on it meaningless.
LP concentration
Named in ../ROADMAP.md as a DEX candidate. Fails Gate 8 outright. The LP
share token is a plain SEP-41 contract (CAVKLYY4… for the XLM/USDC standard pool, wasm
07bab30e…); its complete exported interface is
balance / transfer / transfer_from / approve / allowance / mint / burn / burn_from / decimals / name / symbol / upgrade
— there is no holder enumeration and no holder count. Recovering the distribution needs an
indexer over transfer events, which is the off-chain data path Gate 8 disqualifies however
well-anchored everything else is. Concentrated-liquidity positions are worse: keyed by
(owner, tick_lower, tick_upper) through get_position(...), with no way to enumerate owners at
all.
Price divergence from a reference market
Also named in ../ROADMAP.md. Rejected on the reasoning ROADMAP.md already
records for market-depth-aware oracle scoring: an AMM's price legitimately differs from SDEX by
up to the fee plus whatever arbitrage has not run yet, so a divergence reading cannot separate a
real problem from a normal one. Horizon order books are also trivially spoofed with walls that are
never hit. A cross-pool reference is worse still — it is another market standing in for the one
being scored, which is the substitution Gate 8 exists to refuse.
liquidity_pool_plane as a source for any scored input
The plane (CCABO2IQ…, get(pools) -> Vec) returns type, init args and reserves for many pools in
one call, and is tempting as the bulk read. It is a cache written by the pools
(update(pool, pool_type, init_args, reserves), with ReservesSyncLedger recorded on the pool), so
it can lag the pool it describes. Acceptable for discovery; the pool contract is the source for
anything that reaches a number.
Fee tier as a factor
get_fee_fraction() varies genuinely — standard pools use only 10/30/100 bps, but stable pools
were found at 1, 5, 10, 15, 22, 25, 30 and 50 bps. A higher fee is worse execution, not a failure
mode; grading it would dress a pricing preference as a risk measurement. It belongs inside
depthSafety's round-trip cost, where it already is, and nowhere else.
Reserve imbalance against the pool's target ratio
Real for a stable pool, where drift from 1:1 is a genuine de-peg signal — and meaningless for
standard, where imbalance is the price. A factor only 42 of 340 markets can be graded on is a
per-market rulebook, which ground rule 1 forbids.
Aquarius's own AMM API (amm-api.aqua.network)
Published in their documentation, and it would answer several of the questions above directly. Gate 8. No further discussion.
Concentrated-liquidity-specific factors
Tick distribution and active-liquidity fraction are readable — get_active_liquidity(),
get_slot0(), get_tick(), get_chunk_bitmap_batch() — and genuinely interesting. 26 of 340
pools have them. A factor only those pools can be graded on is a per-market rulebook, same as
reserve imbalance. Revisit only if concentrated pools become the majority of pools.
Question A — resolved: no depth factor until there is a unit of value
Decision: option 4. dex ships with two factors, adminKeySafety and assetControlSafety.
There is no depthSafety, no size floor and no denomination logic, and none may be added
without a further published decision. The problem this resolves, and the three options declined, are
recorded below so none of it is re-proposed from scratch.
It was decided on reversibility, which is the argument that outranked the others. Adding a
factor to a live category later is additive: it lands as a labelled version bump, old scores
stay readable as what they were, and the discontinuity is published rather than hidden. Walking back
a published depth denomination is not symmetric — it would mean revising what stored scores meant,
and this project's own rule is that history is never backfilled across a bump and cannot be,
since risk_scores keeps only outputs and no row can be recomputed. So the two directions carry
very different costs: shipping without depth is a gap that can be closed, while shipping the wrong
depth denomination is a mistake that cannot be undone. Under that asymmetry the thin category wins.
Gate 4 is satisfied by not needing a floor, rather than by having one — and this is a
consequence of the decision, not a dodge. Lending's floor exists because an empty market publishes
0 in the danger band, meaning the opposite of the truth. Neither surviving factor is
size-sensitive: a pool holding 3 stroops has exactly the same seven admin roles and exactly the same
token issuers as one holding 10,000 XLM, and both readings are equally true of it. There is no
quantity here that stops carrying information as a pool gets smaller, so there is nothing for a
floor to protect. A floor becomes necessary again the moment depthSafety lands, which is the
same moment the unit of value to express it in would exist.
The problem itself, unchanged, because it is what any future attempt has to solve:
Aquarius reads no oracle, so nothing in its own contracts values a pool. Two consequences:
- This category's size floor has no denomination. A floor is badly needed: of the 148 XLM-paired pools, 18 hold zero XLM and 59 hold under 1 XLM (many hold 1–3 stroops); only 16 hold more than 10,000 XLM. Over half the registry would otherwise publish a number computed from dust. But "below what?" has no answer in Aquarius's own terms.
depthSafetyis degenerate for 272 of 340 pools if the trade is sized relative to the pool. For a constant-product pool, the price impact of swapping a fixed fraction of the input reserve is a closed form that does not contain the reserves — so everystandardpool at the same fee tier scores identically and the factor collapses into the fee tier it was told not to grade. It would discriminate only amongstable(amplification) andconcentrated(tick distribution) pools. An absolute trade size fixes this, and needs the unit that does not exist.
No number in this document is copied from lending's market-size floor or its minimum-size filter, and none may be: those are denominated in USD against a lending pool's supplied value, which is a quantity this category cannot compute.
The options, with gate verdicts. None is adopted here:
| # | Option | Verdict |
|---|---|---|
| 1 | Denominate in XLM, pool-relative. 148 of 340 pools contain XLM directly, so no price is needed for those. | Clears Gate 8. Leaves 192 pools unscorable and needs a coverage status for them. |
| 2 | Derive prices from Aquarius's own reserve ratios along an XLM-paired path. | Clears Gate 8 on the letter. Almost certainly fails Gates 1 and 2: a price computed by Stenion from a manipulable pool is a fabricated anchor — precisely what lending's rejected "Stenion-computed deviation" candidate already refused. |
| 3 | Anchor to Aquarius's own declared cost of existence — get_standard_pool_payment_amount() = 300,000 AQUA, the price the protocol itself charges to create a pool, and the closest structural analogue to Blend's PoolConfig.min_collateral. | Clears Gate 1 as a type-(a) anchor. Still needs option 1 or 2 to compare a pool's reserves against it. |
| 4 | Score no depth factor at all; publish depth as a route-(a) disclosure and admit dex on adminKeySafety + assetControlSafety. | Honest, but a two-factor category is thin enough that Gate 0 should be re-argued before accepting it. |
Why option 2 was rejected, even though it is fully on-chain
Deriving prices from Aquarius's own reserve ratios along an XLM-paired path clears Gate 8 on the letter — every read is Soroban RPC against the protocol's own contracts, with no aggregator and no off-chain source anywhere. It is rejected anyway, and the reason is worth stating precisely because "it's all on-chain" is exactly what makes it tempting.
A reserve ratio is not a price. It is spot state, and it is cheap to move. A large one-sided swap — or a flash loan, which needs no capital at all — skews a pool's ratio for exactly as long as it takes us to read it. So the "price" would be whatever an adversary wanted it to be at the instant of the read, and the factor built on it would report a comfortable number precisely when someone was active in the pool. A score derived from the same system it is meant to be checking goes blind at the only moment it matters.
This is the same failure shape as the price-deviation candidate already rejected for lending's
oracleSafety (see index.md §2's rejected candidates) — a Stenion-computed figure
standing in for an anchor the protocol does not publish — and it fails two gates rather than one:
- Gate 1: the resulting number is not a protocol-declared anchor. It is Stenion-computed from mutable state, which is the definition of an unanchored threshold wearing an anchor's clothes.
- Gate 2: it does not fail to the unsafe end. It fails by lying confidently — publishing a plausible, precise, wrong number, which is worse than publishing nothing and much worse than publishing 0.
A TWAP-based approach is deferred, not dismissed. Time-averaging would genuinely resist the single-block skew above, and the objection in this section does not defeat it. It is a materially different and harder design — it needs a window, an update cadence, a manipulation-cost model for that window, and a story for pools that trade rarely — and none of those has an anchor in Aquarius's contracts either. Designing it against exactly one protocol would bake Aquarius's shape into a category rulebook. Revisit when a second dex protocol exists to calibrate against.
Why options 1 + 3 were not taken either, despite being the submission's own lean
The original recommendation was XLM-relative sizing plus an absolute floor anchored to the 300,000 AQUA pool-creation cost. It is the strongest of the three declined options and it was still declined, on three grounds:
- It is not reversible the way option 4 is. A published depth denomination becomes what stored scores mean. Changing it later requires revising history, which the no-backfill rule forbids — so the choice would be effectively permanent, made now, on one protocol's data.
- It strands 192 of 340 pools on the category's flagship factor. Only 148 pools contain XLM directly, so the rest would sit in coverage-only limbo — unscored on the very measurement the category was built to publish. A flagship factor that does not apply to 56% of the market is a weak flagship.
- The floor does not answer the question it appears to answer. "It cost 300,000 AQUA to create this pool" is a fact about the protocol's fee schedule, not about whether the pool is deep enough for a number computed from it to mean anything. It clears Gate 1's anchoring requirement — the figure really is protocol-declared — while leaving the actual depth-adequacy judgment unstated and ungraded underneath it. An anchor that anchors the wrong quantity is a worse failure than an admitted judgment call, because it looks anchored.
Both are worth revisiting once a second dex protocol exists to calibrate against. The objection throughout is not that these ideas are bad; it is that every one of them has to be designed against a single protocol today, and a category rulebook designed against one protocol is that protocol's rulebook wearing a category's name.
What holds until then: no adapter may implement depthSafety; no size floor exists, and none is
needed by the two factors that ship; depth is published as a route-(a) disclosure rather than graded.
Comparability: within dex yes, across categories no
Gate 5, stated in both directions, because only one of them is obvious.
Within dex, two safetyScores are comparable. Every protocol scored under this rulebook is
graded by the factors above, with the same formulas and the same thresholds, under the same version
stamp — ground rule 1, which binds every adapter in a category with no exceptions. That is what
makes the ranked block a ranking rather than a list.
Across categories they are not, and nothing may present them as if they were. A dex
safetyScore of 70 and a lending safetyScore of 70 were produced by different factor sets,
different formulas and different weights, from data that has no quantity in common. Neither number
is evidence about the other; "the DEX is safer than the lending market" is not a statement either
score supports.
This is enforced, not merely stated:
buildRegistryView(dashboard/app/lib/registry-query.ts) publishesRankedCategoryGroup[]and no flat ranked array, so each category is its own block numbered 01..n within itself and there is nowhere for a cross-category ranking to live. Name sort is the sole ordering allowed to merge categories, because alphabetical asserts no ranking.- Every entry carries its
categorythrough to the board and to the API responses. risk_scores.categoryis stamped besiderisk_scores.methodology_versionon every run, because both counters start at 1 and the integer alone does not identify a rulebook.
Sharing the adminKeySafety key between the two categories changes none of this. The key names
the question; it does not claim the answers were computed the same way — and the formulas above are
the proof: lending grades one admin account's signer set, this grades seven roles' worst posture
under a pending-upgrade ceiling. See factor 1.
One adapter, many markets
The same rule lending already runs under, and Aquarius has the same shape as Blend: exactly one
wasm per pool type, each matching the hash the router itself declares —
ConstantPoolHash ae0da5a8…de9852 (272/272 pools), StableSwapPoolHash f1077e0b…e747cd (42/42),
ConcentratedPoolHash 12fca5a7…d37ee6 (26/26). Verified against the router's own declared hashes
rather than assumed.
So a second Aquarius market is a config entry and no new scoring code. CLAUDE.md's "one adapter may serve several markets; a market never gets its own adapter" applies here unchanged, and nothing on a per-pool config may be a threshold, weight or formula — that would be a per-pool rulebook.
Operational state is published, never scored
The decision, up front: pause/frozen state is a published field beside the score, and it is deliberately not a factor, not a multiplier, and not any input to a number. It was decided this way on 2026-08-25 after reading both protocols' contracts, and this section exists so it is not re-litigated by the next person who notices a paused pool with an unchanged score.
Both adapters had always read a pause signal — Blend's PoolConfig.status, K2's
router.is_paused() — and neither had ever used it. That could not stay true indefinitely:
every adapter written while it stayed unresolved would have to be retrofitted later.
What each protocol actually means by "paused"
Read from the contracts, not from the documentation — the published docs name Blend's states but publish no numeric mapping, and the mapping circulating in search results is partial and partly wrong.
Blend V2 gates in require_action_allowed (pool/src/pool/pool.rs), which is the entire
rule:
if (status > 1 && (action == 4 || action == 9)) // Borrow, DeleteLiquidationAuction
|| (status > 3 && (action == 2 || action == 0)) // SupplyCollateral, Supply
{ panic!(InvalidPoolStatus) }
RequestType numbering is from pool/src/pool/actions.rs; the setter paths are
execute_set_pool_status (admin) and execute_update_pool_status (permissionless) in
pool/src/pool/status.rs.
status | Blend's name | Borrow | Supply | Withdraw / Repay / liquidation fills | Who can set it |
|---|---|---|---|---|---|
| 0 | Admin Active | yes | yes | yes | admin only (needs the backstop threshold met and Q4W < 50%) |
| 1 | Active | yes | yes | yes | permissionless only (backstop healthy) |
| 2 | Admin On-Ice | no | yes | yes | admin only |
| 3 | On-Ice | no | yes | yes | either — admin, or automatically at Q4W ≥ 30% / below the backstop threshold |
| 4 | Admin Frozen | no | no | yes | admin only; supersedes the backstop, which cannot move it |
| 5 | Frozen | no | no | yes | permissionless only — automatically at Q4W ≥ 60% (≥ 75% from status 2) |
| 6 | Setup | no | no | yes | initialization only; supersedes everything |
Blend never blocks a withdrawal or a repayment at any status. The only user-facing action blocked below the supply threshold is cancelling an in-flight liquidation auction, which is a wind-down-safely posture rather than a restriction on depositors.
K2 (Kinetic) has two layers. storage::is_paused is checked at the top of validate_supply,
validate_withdraw, validate_borrow, validate_repay and validate_liquidation, and again in
the flash-loan and two-step liquidation entry points — so a paused K2 halts everything,
withdrawals included, and deposited capital cannot leave. Separately, each reserve carries its
own gating flags in the ReserveConfiguration bitmap it already publishes its decimals in
(contracts/shared/src/utils.rs, bits 50–53): active, frozen, borrowing_enabled, paused.
A cleared active or a set paused blocks every operation on that reserve; frozen blocks
supplying and borrowing while leaving withdrawals open; a cleared borrowing_enabled blocks only
borrowing.
The shared representation
Because those two vocabularies do not map onto each other, the published state is named by what is blocked, which is the one axis on which the protocols are genuinely comparable:
| Level | Meaning | Blend | K2 |
|---|---|---|---|
active | nothing restricted | status 0, 1 | not paused, every reserve open |
borrowingDisabled | cannot borrow; supply and exit both work | status 2, 3 | reserve borrowing_enabled = false |
entryDisabled | cannot borrow or supply; existing positions can still exit | status 4, 5 | reserve frozen |
exitDisabled | cannot withdraw — capital cannot leave | unreachable | router.is_paused(), or reserve paused/!active |
notOperational | the market was never opened | status 6 | no analogue |
Where a protocol gates per reserve, the most restricted reading is published — the same
worst-reserve convention §2, §4 and §5 use, for the same reason. Alongside the level: the
protocol's own reading verbatim (PoolConfig.status = 4), the exact operations blocked, when it
was read, and whether the value is one only an admin could have set. That last field is
indeterminate for Blend's status 3, because execute_set_pool_status accepts it too — reading
even/odd as "who did this" would be right six times in seven and wrong on the one value where it
matters.
notOperational is reachable in principle and not in practice: every Setup pool in the
2026-08-22 factory survey held exactly $0.00 and is already excluded by
the market-size floor.
Why it is not scored
Three options were weighed — a sixth factor, a multiplier on the overall score, and a published flag. The flag was chosen, on four grounds:
- No on-chain datum resolves the ambiguity. A pause can be an admin containing a threat or an admin abandoning a market, and neither protocol's state carries a reason. Distinguishing them needs off-chain announcements, which adapters may not read. A "context-dependent" factor with nothing to condition on is a flat penalty in costume, and any magnitude for it would be invented — this document's standard is that a threshold is anchored to a protocol's own parameter or labelled an unvalidated judgment call, and there is no anchor here at all.
- The one axis clean enough to score does not exist on both protocols. The strongest scored
variant was not a sixth factor but folding
exitDisabledintoliquiditySafety: that factor is defined as the withdrawal cushion, and a cushion you are contractually barred from drawing on is zero by definition rather than by judgment — no invented magnitude required. It was rejected anyway, becauseexitDisabledis structurally unreachable on Blend. A rule that is live code on one adapter and dead code on the other satisfies ground rule 1 in form only. Recorded here so it is not re-proposed as the obvious fix. - A scored rule would have shipped untested. Every registered market was fully operational
when this landed — Blend status 1, YieldBlox status 0, K2 unpaused with all four reserves open
— so the before/after comparison a scored change requires would have compared each number
against itself. What a factor would have done is worse than nothing: "not paused" scores 100,
so a sixth factor at weight
wraises every active protocol's score by(100 − score) × w, handing Blend, K2 and YieldBlox 5–11 free points for the ordinary state. That is the five-way redistribution the weights note already declined once. - A multiplier would break the score model.
safetyScoreis one published line, and this document's worked example spells the arithmetic out. A client that fetchesfactorsand reproduces the score would stop getting the same number — a verifiability platform whose published factors no longer reconstruct its published score has traded away more than the change buys.
There is also no decay problem, which both scored options have and neither answers: snapping back on unpause puts a step in the history chart that is not a change in risk, and a cooldown invents a time constant from nothing. A published state is a live reading — correct at every instant, with nothing to recover from.
What this costs, stated plainly. A reader who looks only at the number is not protected by the
flag. That is why the flag is a first-class field on both the leaderboard and the detail response
rather than a footnote, and is rendered beside the name and score everywhere either appears — the
same treatment, for the same reason, as deployedOn. If it ever stops being rendered there, the
decision not to score has quietly become a decision to hide.
No version bump. Nothing here changes a formula, a threshold or a weight, and no stored score
moves, so lending's methodology version stays at 1. That the state cannot reach a factor is enforced
rather than intended: adapters/blend/score.test.ts and adapters/kinetic/score.test.ts each
assert a
byte-identical factor map across every restricted state their protocol can be in. If pause state
ever moves a number, those tests fail before the change ships.
Findings are published, not scored — and how they must be written
Verifiable observations we can't or won't grade go in the protocol page's Findings section
(dashboard/app/lib/protocol-notes.ts), never into a factor. Nothing there is read by any
scoring path, and a note — favourable or not — can never move a number.
Findings are the STATIC half of ungraded publication. They are hand-written and reviewed in a PR, which is right for an observation someone had to go and establish, and wrong for a reading that changes every five minutes. The live half is operational state: measured every cycle, published as a typed field, and equally never graded. A new ungraded observation belongs in whichever of the two matches how it is obtained — never in a factor, and never invented as a third mechanism.
A note must survive the history it was drawn from. Twice now a Findings note has outlived
the stored runs behind it: once when the development-era history was discarded, and again when
the briefly-live v2 rows were. Score history is not an archive — it is discarded across a
rulebook change and cannot be recomputed, because risk_scores keeps only outputs. A note
written as "our history shows X" therefore decays into an unverifiable claim on a page whose
entire pitch is that you don't have to trust us.
So every note citing our own observations follows the same form:
- Cite a closed window, with both ends stated. "Between 2026-08-11 18:16 and 2026-08-18 15:55 UTC, 1,469 runs" — not "93% of runs", which silently means something different every time the cron fires. A reader re-running the query later must be able to tell that a different number is a later window, not a contradiction.
- Say the counts are a snapshot of that window and do not update.
- Phrase the underlying claim so it stays checkable from chain after the history is gone. Our runs are evidence that a condition persisted; the condition itself must be one anyone can observe today, directly from the contracts. If the only support for a claim is rows in our database, it is not a finding — it is an assertion.
- Give the exact verification steps — contract, method, field, and what to compare against. If we can't say how a reader would check it themselves, it doesn't go in.
- Claim only what was measured. Where a sub-signal wasn't recorded separately, say so and scope the claim to the runs that carry it, rather than generalising across all of them.
Disputing or changing a threshold
Every number in this document is meant to be challengeable — especially the ones labeled "unvalidated judgment call." If you believe a threshold, weight, or formula is wrong (including if you are a protocol being scored):
- Open a GitHub issue against this repository describing the specific threshold/formula and why you think it's wrong. Anchor your argument to something external where possible (a protocol's own on-chain parameter, a published risk framework, observed data) rather than preference.
- Or open a pull request editing this file directly with the proposed change and its
justification. A change to
methodology/must be accompanied by the matching change to the adapter code (and vice versa) — the two are not allowed to drift. - Maintainer review is required, at the same bar as adapter code changes. A methodology change affects every protocol's number, so it is reviewed at least as carefully as a code change — not merged on preference, and never merged because a scored party requested it. Per the ground rules above, no change is ever accepted in exchange for payment.
Changes that alter what a factor means (e.g. adding or removing a factor) are breaking
changes to the shared taxonomy in core/src/types.ts and are held to a higher bar again —
they affect every adapter at once.