Scoring methodology · v1.0

The whole rubric, in public.

A score you cannot audit is an opinion. This page is the complete method: the arithmetic, what each band obliges you to do, the 7 controls that fail an assessment outright, what this does not cover, and how a provider disputes a result. No part of the score is produced by a language model.

Domains
12
Questions
64
Weight units
20
Automatic fails
7

How the score is computed

How the score is calculated

Each domain is scored independently: answered questions earn 1 point for Yes, 0 for No. N/A answers are excluded from the denominator — they don’t penalise providers for genuinely inapplicable controls.

Domain score = Yes answers ÷ Applicable questions × 100

Overall score = Σ (domain score × weight) ÷ 20

8

High-weight domains

Count 2×

4

Medium-weight domains

Count 1×

20

Total weight units

Denominator

High-weight domains cover controls where failures directly expose participant funds or PII — Access Controls, Encryption, and Business Resiliency among them. 7 specific failures (like no MFA, no encryption, or an unassessed sub-processor) are automatic red flags — see the risk bands to the right for what a red flag does to the final score.

Risk rating bands

The overall score maps to one of four risk bands. Band thresholds are fixed and auditable — no LLM adjusts them at runtime.

  1. Strong

    90–100%

  2. Adequate

    75–89%

  3. Needs Improvement

    60–74%

  4. High Risk

    0–59%

Any single red flag — an automatic-fail control, such as no MFA or unencrypted data — forces the band straight to High Risk, regardless of the numeric score. Red flags are disclosed explicitly in the report so fiduciaries can act on them.

You can’t N/A your way out of a domain

Because N/A answers leave the denominator, marking an entire domain N/A would quietly remove it from the score. Exactly one domain may be excluded that way — Domain 8, the software-development lifecycle, since most plan and fund offices don’t build software. Every other domain must be genuinely answered: a submission that blanks one out is rejected rather than scored, so a weak area can’t be hidden by omission.

What the AI does — and never does

The score above is computed entirely in deterministic code. The AI layer does one thing: it reads the free-text comment behind each “Yes” and grades how well that comment substantiates the claim — a named tool, document, or date scores high; boilerplate scores low; answering Yes with no evidence at all never earns full credit. That produces a second, clearly-labelled evidence-adjusted score shown beside the official one.

It never sets a Yes or No, never moves the official score, and never cancels a red flag. If the AI is unavailable, the assessment still scores normally.

What each band means, and what to do

A band is not a grade for its own sake. Each one carries a prescribed course of action, so the result can go into a board minute with the decision it implies already attached.

Strong90–100%

No material shortfall against the DOL's twelve best practices, and no automatic-fail control missed.

Do this: Retain and re-assess annually. Keep the result in the file as evidence of a prudent selection and monitoring process.

Adequate75–89%

Controls are broadly in place with noteworthy shortfalls in one or more domains — none of them automatic fails.

Do this: Retain, but put the specific gaps in writing to the provider with a remediation date, and re-check at the next cycle.

Needs Improvement60–74%

Shortfalls are significant and concentrated enough that a reasonable fiduciary would not treat this as settled.

Do this: Require a written remediation plan with dates before the next contract renewal. Document the request and the response.

High Risk0–59%

Either the weighted score is below 60, or an automatic-fail control was missed. A single automatic fail lands here regardless of the numeric score.

Do this: Escalate now. Obtain written remediation commitments, consider whether the relationship can continue, and record the decision and its basis.

The 7 controls that fail an assessment outright

Missing any one of these sets the rating to High Risk regardless of the numeric score. They are listed here in full so a provider can check them before submitting, and so a sponsor can see that the escalation is rule-based rather than discretionary.

  1. 1No SOC 2 Type II report or equivalent independent third-party audit (Domain 3).
  2. 2Multi-factor authentication not enforced for remote or administrative access (Domain 5).
  3. 3Participant data not encrypted at rest or in transit (Domain 10).
  4. 4Vendor can move plan assets or process distributions without documented controls (Tiering Q5–Q6 with Domain 5).
  5. 5An undisclosed, or unremediated, past incident (Domain 12).
  6. 6No documented incident-response plan (Domain 9).
  7. 7Sub-processors with access to plan data that are not security-assessed (Domain 6).

What this method does not do

Stated plainly, because a rubric that claims no limits is not credible.

  • This is a self-assessment. Answers are supplied by the service provider; Paladin scores them, it does not independently verify each one. The evidence-quality layer grades how well a written answer substantiates itself, but it cannot confirm a control exists.
  • A score reflects what was true on the submission date. Controls drift; a result more than twelve months old should be treated as stale.
  • Twelve DOL EBSA domains are not the whole of information security. A Strong rating means no material shortfall against this framework, not that the provider cannot be breached.
  • N/A answers are excluded from a domain's denominator. Exactly one domain (software development lifecycle) may be excluded wholesale, and N/A is rejected on automatic-fail controls unless the scoping answers establish it is genuinely inapplicable.
  • The score is not legal advice and does not by itself discharge a fiduciary duty. It is evidence of a documented, repeatable process.

Disputing a result

A provider who believes an answer was recorded incorrectly can contest it. Target response: 5 business days.

  1. 1.The provider, or the plan sponsor on their behalf, submits the question number in dispute and the evidence they say contradicts the recorded answer.
  2. 2.Paladin re-runs the deterministic engine against the corrected answer set. Because the rubric is published and the arithmetic is fixed, the effect of a correction is predictable before it is applied.
  3. 3.If the correction is accepted, the assessment is re-scored and the revised result supersedes the original. Both versions are retained — the original is not deleted, because the sponsor's file needs the history.
  4. 4.If it is declined, the reason is given in writing against the specific rubric line relied on.

Disputes: methodology@paladinassurance.com

Version history

v1.0 — current. Twelve domains, 64 questions, 8 high-weight and 4 medium-weight domains, 7automatic-fail controls. A change to the rubric changes this version, and any assessment records the version it was scored under — so two results are always comparable or explicitly not.

Run an assessment against this rubric