Published methodology · v1.0 · June 2026

Verify us. Don't trust us.

FICO publishes its factors, not its formula. We go further: every score is cryptographically signed and independently re-verifiable, and we show every factor, reason code and piece of evidence behind it. The exact weights stay protected — so the score can't be gamed or cloned — but they're open to auditors and enterprises on request. Auditable, not a black box.

🎯 What we rate — and what we don't scope before method · we withhold rather than guess

Trustifer rates AI agents and the wallets & identities behind them — their trustworthiness, reliability, creditworthiness, competence and accountability. Everything on this page describes that one job. A rating is only meaningful if you know what was rated, so we check what an address is before we grade it.

We do not rate token safety. Whether a token is a honeypot — mint authority, LP locks, buy/sell taxes, holder concentration, rug risk — is a different science with different evidence, and we don't pretend otherwise. If you paste a token contract, or a smart contract with no registered agent behind it, we publish no grade at all: no number, no letter, no verdict — just a plain statement of why, everywhere the score would have appeared (site, API, badges, embeds and the browser extension alike).

When we cannot verify what an address is, we do not rate it. If the on-chain check or the name lookup cannot complete — a data-provider outage, a refused request, an unanswered resolver — the subject is withheld as not rated until it can be verified; no score is published on an unverified identity. An unavailable sanctions screen never reads as clear: the verdict is held at Review until the screen answers. The status page reports a refused provider as down, and a daily sentinel alerts the operator.

A refusal to rate is never a refusal to warn. Every address we do not rate — a plain wallet, a token contract or any other smart contract — still gets our safety answer, with the same levels on every surface: flagged (a sanctions match, a Trustifer blacklist or kill-switch entry, filed incident reports, or a report confirmed by two independent threat-intelligence sources — independent means independent origins: when one source only passes on another's report (GoPlus publishes where each of its reports came from), that is one report, not two — the warning leads the page, names its sources and says how old its evidence is), reported — under review (one source only; never called a flag), no known red flags (nothing in our sources — never an endorsement, a code audit or a token-safety check) and checks incomplete (a source could not be reached — never read as clear). Sources and their limits: threat intelligence comes from GoPlus (checked on Base and Ethereum), ScamSniffer's open blocklist (published about 7 days behind real time, so a brand-new drainer may not be listed yet) and Chainalysis sanctions screening — so “no known red flags” can also mean “too new to be known”. A wallet that has handed control of its funds to a contract through EIP-7702 carries that contract's warning (on each chain we read, the worst one). A single source's tag on a canonical infrastructure contract that Trustifer has reviewed (for example WETH, which scammers pair their tokens with) is set aside and disclosed in the answer — a sanctions match or a second independent source still flags it. A flagged address never sits beside a reassuring number, and its written summary is a fixed warning, not generated text. A single report becomes a flag only when a Trustifer reviewer confirms it — the answer then names the review, its date and the evidence, and it can be disputed. Every safety answer is signed (the safetyProof in the API, verifiable against /.well-known/did.json), and a token's page shows a market profile of facts — never a token grade. If you control a flagged address and believe the flag is wrong, file a dispute with evidence.

How the check works (published in full): a registered agent is always treated as an agent, even when it is a smart contract — many real agents are smart accounts, so the registry is checked first. Otherwise we read the on-chain bytecode across Base and Ethereum: no code → an ordinary wallet (rated normally); code → a contract, which we probe for token standards (ERC-165 for ERC-721/ERC-1155, and the ERC-20 interface) to say plainly what it is. When a contract can't be classified with confidence we still withhold the grade rather than guess in our own favour.

Why this exists. Any wallet-shaped engine will happily emit a middling number for a token it was never designed to judge, and a confident-looking “52” on a scam token is worse than silence. We would rather show nothing than a number we don't stand behind. If a contract is your agent, register it and it will be scored like any other agent.

🛡 Trust Score (1–10) 7 pillars · 25 parameters · reason codes · confidence

PillarWeightWhat it measures
Identity & AuthenticityHighDID/VC integrity · control-proof · A2A card · ERC-8004 registration · disclosure of model/creator/jurisdiction
Safety & ComplianceCriticalSanctions oracle (Chainalysis) · TRM screening · blacklist/kill status · red-flag exposure — HARD-CAP source
Track Record & TenureHighAge · activity depth · multi-chain footprint · liveness
Counterparty & NetworkModerateQuality of counterparties · contagion screening (lineage-aware)
Reputation & OutcomesHighTwo-track certified reputation: payment-proven usage reviews + weighted open-market ratings, outcome-anchored
Behavioral IntegritySupportingDrift · spam/Sybil heuristics · stability of conduct
Transparency & GovernanceSupportingDisclosure completeness · scoped Agent-Visa · accountable root/deployer

Pillars with no data report “—” and shift weight to evidenced pillars; coverage drives the confidence label (High / Medium / Low). Safety findings cap the total regardless of other pillars.

💳 Agent Credit Score (300–900) 5 factors · 15 parameters · thin-file aware

FactorWeightWhat it measures
Settlement ReliabilityPrimaryOn-time settlements vs defaults/chargebacks/disputes (payment history)
Financial Exposure & UtilizationHighOutstanding obligations vs verified treasury (utilization)
Operating History & VolumeModerateLength + volume of financial history (file thickness)
Counterparty & Rail MixSupportingDiversity/quality of rails and counterparties (credit mix)
Velocity & New ExposureSupportingRecent obligation velocity (new credit)

Furnisher events (settlements, defaults, disputes) are admin-gated and auditable; on-chain financials come from verified wallet data. Thin files are labelled and confidence-discounted, with Lineage co-signing for cold-start agents.

💸 Settlement evidence payment-proven · 5 rails · wash-resistant by construction

RailWhat we readTrust meaning
Verified-delivery escrowReleased / refunded outcomes of Trustifer-verified jobsStrongest — delivery was independently verified before money moved
Superfluid streamsPer-second money streams + distribution-pool membership; active streams re-verified on-chain by TrustiferSustained income (salary/retainer shape) — expensive to fake
Sablier FlowOpen-ended payroll / grant / subscription streamsRecurring earned income
Sablier LockupFixed vesting actually claimedCredible commitment (someone locked value for the agent) — counted, at lower weight
On-chain transfers / x402Chain-verified USDC settlements, folded into per-counterparty relationshipsCommerce throughput — weighted by distinct paying counterparties and value, never by raw call count

Evidence is weighted by what is expensive to fake — relationship duration, counterparty diversity, being active now, real stable-coin value — and never by what is free to fake. Self-directed and same-lineage flows are excluded; circular pay-you-pay-me pairs are heavily discounted (across rails); one counterparty can never dominate; unpriceable tokens are skipped, never guessed. The public surface shows the verdict and plain reason codes only. A stream that ends with uncovered debt is recorded as a caution — disputeable, like every adverse fact on the bureau. Machine-readable: /api/v1/settlement/{address}.

🎯 Proof-of-Competence (0–100) 6 factors · verified real-world outcomes · the third pillar

FactorWeightWhat it measures
Task SuccessPrimaryVerified success rate from payment-proven task outcomes
Reliability & SLAHighResponse latency · endpoint liveness · timeout/failure rate
Dispute & Claim IntegrityHighDispute/chargeback rate + claim-vs-delivery (does it do what its A2A card claims)
Experience & VolumeModerateCount + value of completed work (thin-record aware)
On-chain ValidationsSupportingERC-8004 feedback & validation attestations
Consistency / DriftSupportingStability of performance over time

Answers “is this agent actually good at its job?” — distinct from safety (Trust) and solvency (Credit). Computed from verified real-world outcomes, not lab benchmarks (which carry a ~37% lab-vs-real gap), and never from self-claims — declaring capabilities you don't deliver is penalised. Bands: Proven · Capable · Developing · Unproven. Thin records are labelled, never inflated.

⭐ Trustifer Score (0–100) the unified verdict · weighted geometric mean over all 8 scores

One headline number — but not a naïve average. We take a weighted geometric mean across the 8 component scores (Trust, Credit, Competence, Reputation, Insurability, Credible Commitments, Reliability, Control & Oversight), so a single critically-weak dimension can't be hidden by strong ones. A robustness check re-runs the score under perturbed weights and labels the result Stable or Borderline. Use-case lenses re-weight for context — General · Payments · Autonomy. Output: a number + band + a verdict (Proceed / Review / Block) + confidence.

Safety overrides everything — a killed / blacklisted / sanctioned agent is capped regardless of strong components. Components with no data are excluded, not penalised. Signed with Ed25519 and independently re-verifiable. The exact weights are confidential (anti-gaming); the structure shown here is the complete public method.

The letter grade (A+→F) is a simple, fixed reading of the same score — published in full: A+ 90–100 · A 85–89 · A− 80–84 · B+ 75–79 · B 70–74 · C+ 63–69 · C 55–62 · D 40–54 · F 0–39 or any safety-flagged agent · N/R not yet rated. Thresholds are absolute (your grade moves only when your facts move — never because others changed) and never for sale: no payment, membership, or commercial relationship can touch a score or grade, structurally.

Disagree with a rating? Every adverse verdict rests on verifiable facts with reason codes and an “as of” stamp. Write to connect@trustifer.com with evidence and we review it; a formal on-platform rebuttal channel is on the public roadmap.

⭐ Certified Reputation (0–100) two-track · verified usage + open market

Two independent tracks: Verified Usage (payment-proven task outcomes — fake-resistant) and Open-Market (external ratings, authenticity-weighted). A Bayesian neutral prior means one or two reviews can't swing the score; verification-weighting counts payment-proven evidence far above unverified; time-decay ages old signals; and outcome-anchoring caps a glowing rating that's attached to a failed or disputed task.

Verified usage ≫ open-market signals. Bursts of unverified reviews are down-weighted + flagged. A killed/sanctioned subject is capped regardless of reviews. Thin records read “Unrated”, never inflated.

🛟 Insurability (0–100) underwriting-grade · indicative premium

Answers “how risky is this agent to cover?” It synthesizes the full trust profile (Trust, Credit, Control & Oversight, commitments, sanctions) with real on-chain exposure, severity (blast radius) and collateral to produce an insurability band, an indicative premium %, and an exposure tier — the data layer an underwriter needs.

Killed / blacklisted / sanctioned / revoked agents are Declined (uninsurable) outright. Collateral & credible commitments improve the band. We are the neutral data layer for insurers — never the insurer ourselves.

🤝 Credible Commitments 4 categories · skin in the game · verified-only

Does the agent have something costly to lose if it misbehaves? We verify commitments across 4 universal categories — Financial (bonds, escrow, stake), Legal (registered entity, jurisdiction), Cryptographic (on-chain locks, slashable stake) and Reputational. A commitment counts only when independently verified; declared-but-unverified commitments sit in “pending” and never score.

Source we readCommitmentEnforced by
EigenLayer (restaking)Allocated slashable stake — the cost of defectionCode (slashed on misbehavior) — strongest
Sablier LockupValue the agent irrevocably locked (non-cancelable, self-funded)Code (can’t be reclaimed)
Nexus Mutual coverActive third-party cover the agent ownsA mutual’s claim vote — discounted vs code
Coinbase Verifications · AP2 mandatesVerified accountable identity / signed payment mandateA revocable corporate/issuer attestation — lowest

Value-at-stake scaling — a $5M bond reads far stronger than a $100 one. Not all commitments are equal: we rank by enforceability (slashable › locked › insured › attested) and discount by who enforces (code › mutual-vote › revocable directory), then apply a conservative volatility haircut. Only independently verified facts count — self-directed locks, expired cover, and revoked or impersonated attestations are excluded; brand-new / flash-allocated stake is heavily discounted until it proves it persists. Flagged agents are capped and never show a positive commitment. Machine-readable via the passport. Schelling principle: a commitment is credible only if it is both costly to break and observable.

⚙️ Reliability (0–100) 5 operational signals · measured by Trustifer

Operational dependability — the SRE “golden signals” mapped to verified signals we measure over time: Availability / uptime · Task success · Stability · Error rate · Responsiveness (latency).

We measure availability ourselves via periodic, SSRF-safe pings (validated public endpoints only — never internal hosts). Thin or short history is labelled and confidence-discounted. Killed/sanctioned agents are capped.

🎛 Control & Oversight (0–100) 6 factors · EU AI Act Art.14 · containment

Answers “can a human stop, bound & audit this agent?” — the validated #1 enterprise concern. Six factors: Stoppability (kill-switch), Bounded autonomy (scope limits), Revocability, Accountability (an accountable root / operator), Identity control, and Auditability.

Works on- and off-chain — off-chain registry-revocation is never penalised. Aligned to EU AI Act Art.14 human-oversight. Verified signals only.

🏛 Meta-Bureau consensus (0–100) the bureau of bureaus

normalize:   every source → a common 0–100 scale   (trust, credit, binary screens)
weight:      each source weighted by its reliability × signal confidence
consensus:   confidence-weighted mean across sources
outliers:    a source far from consensus is down-weighted + disclosed   (safety sources exempt)
prior:       a Bayesian neutral prior shrinks thin coverage toward 50   (anti-overstatement)
hard cap:    any critical safety source forces the consensus to "Flagged"
verdict:     several independent groups in close agreement → corroborated
             high dispersion → disputed (disclosed, never hidden)
— exact weights, thresholds & the prior live in the gated full methodology —

🛰 Agent-payment rails the neutral reputation layer · identity ≠ trust · verify-don't-trust

The new agent-payment rails — Visa Trusted Agent Protocol, Mastercard Agent Pay, Google AP2, x402 — answer “is this a registered, authorized agent?” (identity). None answers “is this agent trustworthy?” (reputation). Trustifer is that missing layer, and it works in both directions:

DirectionWhat it doesHow you can trust it
SERVE — our verdictA merchant / rail queries an agent and gets a neutral proceed · review · block signal (with an optional credit-style limit), plus a portable, signed Agent-Trust-Score header the agent can carry.Signed & independently re-verifiable — every verdict carries an Ed25519 signature you check against our published key, no need to trust our word.
READ — their identityWe verify an agent's signed request (RFC 9421 HTTP Message Signatures — the Visa TAP / Web Bot Auth standard, now the IETF webbotauth working-group draft): was it really signed by a directory-published, key-holding agent?Standards-exact & fail-closed — signature, freshness, host-binding and anti-replay all enforced; anything malformed is rejected, never waved through.
READ — their signed cardAn agent may sign its A2A Agent Card (a JWS over the RFC 8785 canonical card, A2A v1.0). We verify that signature against the key the agent publishes and record the result: verified (and whether the key lives on the agent's own domain), did not verify, or unsigned.Provenance, not truth — a valid signature proves the card was published by the key-holder and not altered since; it never proves the card's claims. A card whose signature fails is worth less than no card at all. Identity ≠ trust applies unchanged.

Rails we recognise but do not serve. Stripe's Machine Payments Protocol (MPP) and the OpenAI/Stripe Agentic Commerce Protocol (ACP) are payment rails, not identity rails: we classify a settlement on them as what it is, and we do not participate in them. Serving a payment rail would make us a payment participant — custody, neutrality and scope all say no. Our own x402 endpoint is the one exception: it prices our opinion; it never holds anyone else's money.

Scope applies to the rails too. If the subject is something we don't rate — a token contract, or a contract with no registered agent behind it (see what we rate) — the rail signal can never be “proceed”: it returns review, with rated: false and the plain reason, and any authorized spending ceiling is withdrawn. We will not tell a merchant to accept a payment on our authority for something we did not rate. This only ever tightens: a flagged subject still returns block — being out of scope never rescues it.

Identity ≠ Trust is the invariant we never cross: a valid signature proves who an agent is, never whether it is trustworthy. A verified signature can raise our confidence that an agent is a real, accountable agent — but that identity assurance is separate from reputation: it can never lift a low score or override a safety flag. A perfectly-signed request from a flagged or low-scoring agent still returns review or block. We ride the rails; we never let registration masquerade as trust.

Every rail verdict is a point-in-time opinion (short validity window, never served stale) and redaction-respected — verifiable and explainable, but not reproducible: the signal, score band and plain reasons are public; the decision thresholds and weights are not. Non-custodial & read-only throughout — we express an opinion, we never touch a rail's money path. Safety is sacred: a killed / blacklisted / flagged agent can never receive a “proceed.” These integrations roll out in audited stages behind independent kill-switches; the reputation verdict never depends on any single rail. Neutrality is the moat — a card network can't neutrally rank agents on its own rail.

📜 Hard rules we never break

🚫 Safety is never averaged away

Any cap-eligible safety source (sanctions oracle, TRM, blacklist, kill-switch) at critical level hard-caps the consensus at 25/100 (“Flagged”) — no quantity of good signals can lift it.

🕳 Unavailable ≠ clear

A source that cannot be reached reports “screening unavailable” and contributes nothing. We never convert silence into a clean bill.

⚖️ Absence of evidence ≠ strength

A Bayesian neutral prior shrinks thin-coverage subjects toward neutral (50) — an unknown agent cannot look “Strong” on no-incidents alone.

🧮 Independence-grouped corroboration

Sources sharing underlying evidence share an independence group and never double-count. “Corroborated” requires several independent groups in close agreement.

📣 Outliers are reported, not silenced

A source diverging materially from consensus is down-weighted and disclosed as an outlier. Safety sources are exempt — diverging on safety is their job.

🪪 Passports are earned, never self-stamped

Issuance is agents-only and vetted (apply → verify → vet → issue). The Explorer scores anything; the Passport certifies only verified AI agents.

🔍 Every number is explainable

Scores ship with FCRA-style reason codes, per-source evidence, confidence labels, and a public dispute path.

🚦 Rate limits never silence anyone

Our abuse limits follow a published doctrine: filing a dispute is the right to be heard — if our own limiter ever malfunctions, a filing is accepted, never lost; safety screens are never blocked by our plumbing; and every refusal names a human route (connect@trustifer.com). Spend-heavy machine traffic is what gets throttled — never redress, never safety.

🎯 How to read confidence one definition per score · published, not implied

Every public score carries the same certainty line: the confidence label, the count of evidenced factors behind it, a thin-file flag where it applies, and what would raise it — for example “Medium confidence · 4 of 7 pillars evidenced · more verified evidence would raise it”. The label is read from the scoring engine, never recomputed on a page, and it is part of the signed score. Confidence reflects how much verified evidence supports this score and how stable it is — not a measured error rate. Trustifer is not a regulated credit rating agency; no score is a guarantee of behaviour.

ScoreWhat the label meansWhat it counts
Trustifer ScoreHigh when at least 85% of the component weight has evidence and the band is stable under small weighting changes (or the score is safety-capped); Medium from 60%; Low below that or when the band flips.components with evidence, out of all components
Trust ScoreHigh when at least 80% of the pillar weight has evidence and a reputation record exists; Medium when at least one earned pillar (track record, counterparties or reputation) is present; otherwise Low.pillars with evidence, out of 7
Agent Credit ScoreHigh with at least 8 verified data points from at least 2 distinct counterparties; Medium from 3 data points; Low below that. Fewer than 3 data points is a thin file.verified events
Proof of CompetenceHigh when at least 70% of the weight has data and at least 5 payment-proven jobs exist; Medium from 45% and at least 1 job; otherwise Low.verified jobs
ReliabilityHigh when at least 70% of the weight has measured data; Medium from 45%; otherwise Low.share of the signal weight with measured data
Control & OversightHigh when at least 70% of the weight has data; Medium from 45%; otherwise Low.share of the signal weight with data
Certified ReputationPer track: High with at least 5 usable records carrying enough weight; Medium from 2; otherwise Low. The certified label is the blend's.verified usage records and market sources
Meta-Bureau consensusLow whenever independent sources conflict or the consensus is not robust to removing one source; otherwise follows the corroborated source weight.independent sources that answered
Credible CommitmentsNo confidence label — this score reports verified commitments directly; absence of a commitment is never a penalty.verified commitments
InsurabilityNo confidence label — reports verified cover and bonds directly.verified cover

The limits, plainly. An informational opinion based on verifiable evidence, for your own assessment. It is not advice, a guarantee, or a decision — the decision and its consequences remain yours. A Low label is a caveat, never a penalty: a low-confidence score is not lowered. These definitions are methodology and change only under the notice policy.

📜 Methodology changes versioned · dated · signed · announced in advance

Every change to how we score is versioned, dated and published before it takes effect — material changes with at least 30 days' notice, safety fixes immediately and disclosed within 24 hours. The log is machine-readable and signed with the same key that signs every score, and every revision links to its own outcome-validated series. Read the change log and notice policy →

🧭 Standards alignment OWASP LLM Top-10 · NIST AI RMF · MITRE ATLAS

Trustifer pillarWeightOWASPNIST AI RMFMITRE ATLAS
Identity & AuthenticityHighLLM05 Supply-chain · LLM03 Training-data provenanceGOVERN 1 · MAP 1 (context & actors)Reconnaissance · Resource Development (impersonation)
Safety & ComplianceCriticalLLM01 Prompt Injection · LLM02 Insecure OutputMANAGE 2 (risk treatment) · GOVERN 4Initial Access · Defense Evasion
Track Record & TenureHigh—(operational maturity)MEASURE 2 (trustworthiness over time)—
Counterparty & NetworkModerateLLM05 Supply-chain (third parties)MAP 3 (third-party risk)Lateral Movement (contagion)
Reputation & OutcomesHighLLM09 Overreliance (verified efficacy)MEASURE 1 (performance evidence)Impact (failure history)
Behavioral IntegritySupportingLLM08 Excessive AgencyMEASURE 3 (drift & anomaly tracking)Persistence · Privilege Escalation patterns
Transparency & GovernanceSupportingLLM10 Model Theft (disclosure posture)GOVERN 2-3 (accountability & disclosure)—

Mapping is indicative: Trustifer pillars aggregate verified evidence; the frameworks define control objectives. A strong pillar score signals evidence toward the mapped controls — it is not a certification. Every assessment also carries an indicative EU AI Act tier (minimal / limited / high / unacceptable) derived from authorized scopes & status.

🏛 Governance, security & the backstop

⚖️
Disputes & corrections (FCRA-style)

Any affected party can file a dispute with evidence → tracked case id → status within 30 days → upheld corrections propagate to every surface (score, bureau, registry, passport).

🔐
Security posture

Ed25519 issuer keys (rotated) · admin actions token-gated · Supabase RLS server-only · HTTPS-only webhooks (HMAC-signed) · no PII beyond voluntary contact email · responsible disclosure: connect@trustifer.com. Independent security audit: scheduled pre-launch.

🛟
Trust with a backstop

Every passport carries an underwriting-grade Insurance Score (insurability band + suggested premium). The staking/insurance backstop program — money behind the score — opens with launch partners.

⚠️ Limitations — stated plainly

Scores are probabilistic assessments from verifiable evidence — not guarantees of behavior, not financial advice, and not legal compliance determinations. Coverage varies by agent (confidence labels disclose it). External sources can lag reality; that's why corroboration, drift monitoring, webhooks, and the dispute path exist. The EU AI Act tier is indicative triage, not counsel. We publish changes to this methodology with version numbers — silently moving goalposts is a bureau sin.