Decide what you are reviewing
Before giving an implementation a rating, state what you reviewed. Include the website or app, user and account groups, browsers, operating systems, credential managers, countries and time period. Also list anything that was excluded or not tested. When comparing providers, review them under the same conditions. If a provider does not support a relevant platform or flow, show this as a gap. UseN/A only when a requirement does not apply to the reviewed implementation.
Also state whether each capability is built in, can be configured, requires integration work or requires custom development. This separates the quality of the result from the effort needed to achieve it.
How to read the flow metadata
Expected adoption impact is a qualitative estimate for the full audience. It is not a measure of overall importance and does not affect whether a criterion passes. For example, credential management and Signal API reconciliation have low expected adoption impact but high reliability and lifecycle value.
Grade the criteria and flows
Test important platforms separately. A feature that works on Chrome but fails on Safari should show one pass and one fail rather than one combined result.
Core criteria cover behavior that must work, including security, account binding, completion, cancellation and fallback. Quality criteria cover measurement, testing and maintainability.
Use a simple review table:
Also state how strong the available evidence is:
Rate adoption potential
Do not rate adoption by counting passed criteria. A login strategy, an optional enhancement and a management feature have different purposes. Instead, review four separate areas:
The following table shows which flows contribute to each area. A primary flow provides the main result. A secondary flow improves or protects that result but cannot satisfy the area by itself.
Rate each included website or app separately. If an implementation uses a different flow to achieve the same result, assess it against the same acceptance criteria.
If an area does not reach Basic or was not reviewed, the overall result is Incomplete. Otherwise, give each area the highest level it fully meets.
Then determine adoption potential without averaging the pillars:
Determine adoption potential for web and native separately, then use the lowest result across the included applications, ordered Incomplete < Low < Mid < High. An application that was not reviewed counts as Incomplete. Always show the individual results and state what was not included. For example: Incomplete overall; web: High; native: Incomplete (not reviewed).
This rates adoption potential, not actual adoption. A live deployment should also report enrollment, passkey login share, success, fallback and error rates for a stated user group and time period.
For automatic creation, calculate the upgrade rate as successful server registrations divided by conditional-registration requests sent after a completed password sign-in. Do not estimate whether the credential provider considered a request eligible.