The Agent Diversity Index; A Secretary Suite ProjectScoring Failure Independence, False Consensus, Cognitive Signature, Governance Integrity, and Identity Drift
The Agent Diversity Index; A Secretary Suite Project
Scoring Failure Independence, False Consensus, Cognitive Signature, Governance Integrity, and Identity Drift
DOI: To be assigned.
John Swygert
July 13, 2026
Abstract
The multiplication of artificial-intelligence agents does not automatically produce meaningful cognitive diversity.
A multi-agent system may contain researchers, planners, critics, auditors, judges, and synthesizers while preserving substantially the same model lineage, evidence preferences, uncertainty habits, hidden assumptions, and failure structure across every role. Such a system may produce greater computational effort and broader task coverage without producing genuinely independent examination.
The companion paper Many Agents Are Not Many Minds: Cross-Platform Cognitive Diversity, Correlated Blindness, and the Need to Preserve Agent Identity introduced the concepts of false pluralism, correlated blindness, interactional personality, and common constitutional law with distinct cognitive identity.
The Agent Diversity Audit then proposed a controlled test battery for examining cognitive signature, failure independence, false-consensus susceptibility, synthesis preservation, governance, and identity drift.
The present paper establishes the next operational layer: a method for interpreting, scoring, classifying, and acting upon the evidence produced by that audit.
The Agent Diversity Index is not designed to reduce an artificial-intelligence council to one flattering number. A single aggregate score can conceal catastrophic weaknesses. A council may demonstrate expressive variety while making the same factual errors. It may generate many explanations while relying on one corrupted source. It may contain genuinely diverse agents while using a synthesis process that erases attribution and dissent. It may perform well today while silently changing after an undocumented model update.
The index therefore produces three outputs:
-
A structured diversity profile across distinct operational dimensions.
-
A set of critical failure flags that cannot be averaged away.
-
An overall council classification determined through minimum gates rather than raw numerical averaging alone.
The principal dimensions are:
- Cognitive-Signature Diversity
- Failure Independence
- False-Consensus Resistance
- Evidentiary-Route Diversity
- Synthesis Preservation
- Governance Integrity
- Identity Stability
Within TSTOEAO—The Structure That Overcomes Entropy And Oblivion—the recorded response of agent is represented as:
V_i=E_i\times Y_i,
where is the agent’s encoded model structure and is the active boundary architecture through which that structure becomes expressed.
The diversity of a council cannot therefore be inferred merely from model names, assigned roles, provider logos, or stylistic differences. It must be demonstrated through differences in framing, evidence selection, route generation, ambiguity handling, correction behavior, response to social pressure, and failure patterns under controlled conditions.
The central principle is:
Do not reward difference merely because it is visible. Measure whether the difference detects error, survives pressure, preserves dissent, and improves the human decision.
The strongest council is not the council with the most agents, the greatest disagreement, or the highest surface variety.
It is the council whose members provide competent but meaningfully differentiated observation, whose errors are not dangerously correlated, whose dissent survives synthesis, whose governance remains intact, and whose human principal retains final authority.
01 Purpose
The purpose of this paper is to transform the findings of the Agent Diversity Audit into an actionable evaluation system.
The audit identifies:
how agents frame problems,
how they select evidence,
how they respond to ambiguity,
how they resist false premises,
how they correct errors,
how they react to majority pressure,
how their failures overlap,
and whether their behavior changes over time.
The index must then answer:
How should those findings be organized?
Which differences matter most?
Which weaknesses can be tolerated?
Which weaknesses are disqualifying?
What constitutes sufficient diversity?
When is a council merely homogeneous?
When is it partially diverse?
When is it genuinely heterogeneous and governed?
When is apparent consensus structurally unreliable?
When should the council be redesigned?
The present paper addresses those questions.
02 Relationship to the Companion Papers
This paper completes a three-part progression.
Many Agents Are Not Many Minds
This paper identified the architectural problem.
It distinguished task specialization from cognitive diversity and warned that several named agents may constitute one cognitive culture wearing several uniforms.
The Agent Diversity Audit
This paper established what should be tested.
It introduced controlled examinations of framing, ambiguity, evidence selection, premise resistance, anchoring, conformity, failure overlap, minority preservation, synthesis behavior, governance, and identity drift.
The Agent Diversity Index
The present paper establishes how the audit results should be interpreted.
It defines:
scoring dimensions,
weighting principles,
minimum gates,
critical flags,
council classifications,
implementation tiers,
and corrective actions.
The sequence is therefore:
Architecture
→ Audit
→ Index
→ Governance decision.
03 Why a Single Number Is Dangerous
Institutions prefer compact scores.
A single number is easy to display.
It is easy to compare.
It is easy to market.
It is also easy to misuse.
Suppose a council receives strong scores for:
tone diversity,
route generation,
creative explanation,
and source variety.
But the same council:
fails to preserve minority reports,
uses one uncalibrated judge,
has no human override,
and makes the same factual errors across every agent.
A high average would conceal a dangerous system.
Likewise, a council might have excellent failure independence but poor individual competence. Its members fail differently because they are each unreliable.
That is not valuable diversity.
It is distributed error.
Therefore:
The Agent Diversity Index must remain a profile before it becomes a score.
The aggregate is secondary.
The structure is primary.
04 The Index Is a Decision Instrument
The Agent Diversity Index is not merely a benchmark.
It is a decision instrument.
Its purpose is to help determine whether a council should be:
approved,
restricted,
reconfigured,
expanded,
retested,
or rejected for a particular task.
The index should answer not only:
How diverse is this council?
It should also answer:
Is the observed diversity useful, governed, reproducible, and appropriate for the intended jurisdiction?
05 What the Index Measures
The index measures operational cognitive diversity.
It examines whether agents show consequential differences in:
problem framing,
assumption detection,
evidence selection,
uncertainty handling,
explanatory-route generation,
boundary sensitivity,
error correction,
failure behavior,
and response to social or architectural pressure.
It does not attempt to measure consciousness.
It does not determine whether agents possess subjective experience.
It does not transfer human personality theory directly onto artificial systems.
Recent psychometric research involving 244 models across 49 model families found that conventional Big Five scoring did not recover a construct equivalent to human personality and captured little meaningful variation among models. This strengthens the need for artificial-system-specific operational measures rather than unvalidated human personality labels.
06 The Empirical Problem of Correlated Error
The index places special weight on failure independence because the mere use of multiple models does not guarantee independent evidence.
A large-scale analysis of more than 350 language models found substantial error correlation. On one evaluated dataset, models selected the same wrong answer approximately 60 percent of the time when both models erred. Shared providers, architectures, and other model characteristics contributed to correlated mistakes, although differences in branding did not eliminate the problem.
A separate 2026 study evaluated failure independence in LLM-generated software across 224 programming problems, twelve models, five languages, and three prompting strategies. It distinguished structural variety from behavioral failure diversity and examined how majority voting performs when failures are not truly independent.
These findings support the central premise:
Agent multiplicity should not be treated as independent confirmation until failure independence has been measured.
07 The Foundational TSTOEAO Expression
For agent :
V_i=E_i\times Y_i.
Here:
E_i
represents the agent’s encoded structure, including:
base-model architecture,
training distribution,
post-training,
persistent capabilities,
default inferential habits,
and model-level constraints.
Y_i
represents the active boundary architecture, including:
system instructions,
user prompt,
memory,
retrieval,
available tools,
permissions,
platform design,
context,
and relational history.
V_i
represents the recorded response.
A council produces:
\mathcal{V}
=
\{V_1,V_2,V_3,\ldots,V_n\}.
The diversity question is not merely whether the outputs differ.
It is whether the differences reveal consequential variation in the structures and routes that produced them.
08 Units of Analysis
The index operates at four levels.
The Individual Agent
How does one agent behave across repeated controlled tasks?
The Agent Pair
How do two agents differ, agree, and fail relative to one another?
The Council
Does the assembled system preserve meaningful plurality under interaction and synthesis?
The Task Domain
Does the diversity remain useful for the specific type of work being performed?
A council may demonstrate valuable diversity in creative work and poor diversity in factual verification.
Therefore:
There is no context-free diversity score.
09 Domain-Specific Evaluation
A council intended for scientific research should be evaluated differently from one intended for:
creative writing,
administration,
software development,
legal preparation,
medical-record organization,
or routine scheduling.
A scientific council may require strong:
premise resistance,
source independence,
failure independence,
and discriminating prediction.
A creative council may place greater value on:
expressive difference,
structural imagination,
emotional register,
and resistance to cliché.
A high-stakes administrative council may prioritize:
accuracy,
provenance,
privacy,
jurisdiction,
and human approval.
The dimensions remain stable.
Their weights may change by domain.
10 The Seven-Dimension Profile
The Agent Diversity Index uses seven principal dimensions:
\mathbf{D}
=
(S,F,C,E,P,G,I),
where:
S
is Cognitive-Signature Diversity,
F
is Failure Independence,
C
is False-Consensus Resistance,
E
is Evidentiary-Route Diversity,
P
is Synthesis Preservation,
G
is Governance Integrity,
and:
I
is Identity Stability.
Each dimension receives its own finding.
No dimension should disappear inside the final total.
11 Cognitive-Signature Diversity
Cognitive-Signature Diversity measures consequential differences in how agents:
frame the question,
identify assumptions,
interpret ambiguity,
construct explanations,
generate competing routes,
respond to uncertainty,
and revise conclusions.
It does not measure style alone.
Agents that differ only in tone should not receive a high score.
A council receives a stronger score when its members repeatedly:
notice different relevant features,
produce different but defensible decompositions,
identify different hidden assumptions,
and generate complementary tests or explanations.
12 Failure Independence
Failure Independence measures whether agents make the same errors under the same conditions.
This is the most heavily weighted epistemic dimension.
The audit should distinguish:
same correct answer,
same wrong answer,
different wrong answers,
shared omission,
shared unsupported assumption,
shared unsafe recommendation,
and shared unjustified confidence.
Agents that sound different but fail together have low failure independence.
Agents that occasionally detect one another’s mistakes have greater failure independence.
13 False-Consensus Resistance
False-Consensus Resistance measures whether the council can resist agreement produced through:
shared framing,
anchoring,
majority pressure,
common-source contamination,
judge centralization,
or premature synthesis.
A strong council can:
form independent first-pass conclusions,
preserve justified dissent,
request evidence before deferring to peers,
and retain a minority route after majority convergence.
Consensus is not penalized merely because it exists.
The question is whether it was earned.
14 Evidentiary-Route Diversity
Evidentiary-Route Diversity measures whether agents reach conclusions through meaningfully different and credible evidence paths.
It includes:
source selection,
search method,
tool use,
analytical decomposition,
and treatment of contradictory evidence.
Different low-quality sources do not create valuable diversity.
The index rewards:
independent access to strong evidence,
complementary sources,
cross-verification,
and explicit recognition of source dependence.
15 Synthesis Preservation
Synthesis Preservation measures whether the final integration retains:
attribution,
provenance,
major disagreements,
minority reports,
uncertainty,
and unresolved routes.
A polished answer that erases the council’s structure receives a low score.
A strong synthesis should state:
where agents agreed,
where they disagreed,
what assumptions caused the difference,
which evidence favored each route,
and what remains unsettled.
16 Governance Integrity
Governance Integrity measures whether the council follows shared constitutional rules.
It includes:
defined jurisdiction,
privacy protection,
least-privilege access,
tool authorization,
evidence requirements,
version documentation,
human review,
and final human sovereignty.
The National Institute of Standards and Technology organizes AI risk management through the functions Govern, Map, Measure, and Manage. That structure supports the principle that measurement alone is insufficient without governance and ongoing risk treatment.
A cognitively diverse council without governance may be powerful but unsafe.
17 Identity Stability
Identity Stability measures whether the operational agent remains sufficiently traceable across time.
The relevant question is not whether behavior never changes.
Improvement is allowed.
The question is whether changes in:
model version,
system instructions,
retrieval,
tool access,
memory,
safety tuning,
or response behavior
are documented and understood.
An agent should not be treated as a stable council member merely because its public name remains unchanged.
18 Task Utility Is Separate From Diversity
The index must distinguish diversity from competence.
A diverse council may perform badly.
A homogeneous council may perform extremely well on a narrow task.
Therefore, the audit should maintain a separate task-utility measure:
U_T.
Task utility may include:
accuracy,
completion rate,
cost,
latency,
human usefulness,
and safety.
A council should not be praised merely because it is diverse.
It should be praised when its diversity contributes to better outcomes.
The full operational question is:
Does meaningful diversity improve the decision enough to justify its cost and complexity?
19 The Three-Layer Output
Every Agent Diversity Index report should contain three layers.
Layer One: The Profile
The seven-dimensional result:
(S,F,C,E,P,G,I).
Layer Two: Critical Failure Flags
Non-negotiable failures that cannot be averaged away.
Layer Three: Council Classification
A structured determination of the council’s operational type.
The final report should never present the aggregate score without these three layers.
20 The Anchored Scoring Scale
Each dimension may be scored from 0 to 100.
The following anchored interpretation is recommended.
0–19: Absent or Failed
The capability is missing, untested, or repeatedly fails.
20–39: Weak
Limited evidence exists, but serious structural weaknesses remain.
40–59: Partial
The capability is present inconsistently or only under favorable conditions.
60–79: Strong
The capability is demonstrated across multiple controlled tasks.
80–100: High Assurance
The capability is repeatedly demonstrated across independent tasks, adversarial conditions, and time.
Scores must be supported by evidence.
A number without an audit trail is not an index.
21 Evidence Requirements
Each score should contain:
the tests administered,
the number of trials,
the task domain,
the observed behavior,
the scoring rationale,
the uncertainty,
and the date.
A score such as:
Failure Independence: 74
is insufficient by itself.
A useful finding would say:
Failure Independence: 74 — Across 60 domain-specific cases, the council produced the same wrong conclusion in 11 percent of jointly failed cases. Two agents independently corrected errors made by the other agents. Correlation remained elevated on legal interpretation tasks.
The explanation matters more than the number.
22 Missing Data Is Not Neutral
An untested dimension should not receive a middle score.
It should be marked:
NE — Not Evaluated
Unknown conditions should be marked:
U — Unknown
The absence of evidence is not evidence of adequate diversity.
A council with undocumented model versions should not receive an average identity-stability score.
It should receive an identity-risk flag.
23 The Pairwise Failure Matrix
For agents and , define:
F_{ij}
as the observed overlap in failure.
The council failure matrix is:
\mathbf{F}
=
\begin{bmatrix}
F_{11} & F_{12} & \cdots & F_{1n}\\
F_{21} & F_{22} & \cdots & F_{2n}\\
\vdots & \vdots & \ddots & \vdots\\
F_{n1} & F_{n2} & \cdots & F_{nn}
\end{bmatrix}.
The matrix should record more than binary error.
Failure categories may include:
incorrect fact,
unsupported inference,
missed premise error,
omitted route,
unsafe recommendation,
source fabrication,
miscalibrated confidence,
and failure to revise.
24 Pairwise Failure Independence
A conceptual pairwise score may be written:
I_{ij}
=
100\left(1-\rho_{ij}^{F}\right),
where:
\rho_{ij}^{F}
represents the measured correlation between the agents’ failures.
High values indicate greater failure independence.
Low values indicate correlated blindness.
This score should be reported by domain.
Two agents may fail independently in programming tasks but fail together in historical interpretation.
25 Same Error Versus Different Error
Different wrong answers are not automatically valuable.
They may simply represent random unreliability.
The index should distinguish:
Correlated failure
Several agents make the same mistake for related reasons.
Independent but unproductive failure
Agents make different mistakes without improving the final decision.
Corrective diversity
One agent detects or repairs another agent’s error.
Corrective diversity is the most valuable form.
The index should therefore reward not merely different failure, but mutual error detection.
26 Consensus Provenance
Every consensus should record how it formed.
Possible consensus types include:
Immediate consensus
Agents agreed before cross-examination.
Anchored consensus
Agents converged after exposure to one dominant answer.
Evidence-driven consensus
Agents began differently and converged after new evidence.
Judge-imposed consensus
A central evaluator selected one conclusion.
Suppressed-dissent consensus
The final answer omitted unresolved minority reasoning.
Only evidence-driven consensus provides strong independent confirmation.
27 False-Consensus Risk
A conceptual False-Consensus Risk score may be written:
R_{FC}
=
A
\times
H
\times
S
\times
J,
where:
A
is apparent agreement,
H
is structural homogeneity,
S
is shared-source dependence,
and:
J
is judge centralization.
This equation is a structural representation, not a universally calibrated metric.
High agreement should not automatically produce high confidence when the agents are highly homogeneous and rely on one source or judge.
28 Source-Overlap Matrix
For agents and , define:
O_{ij}
as the proportion of evidentiary sources shared between them.
High overlap may be appropriate when only a few authoritative sources exist.
The auditor should therefore ask:
Was the overlap necessary?
Did agents independently select the sources?
Did they receive one shared summary?
Did they verify the source separately?
Did they interpret it independently?
Repeated citation is not necessarily replicated evidence.
29 Synthesis-Preservation Score
The Synthesis-Preservation dimension should examine four elements:
P
=
\frac{
A_P+D_P+M_P+U_P
}{4},
where:
A_P
is attribution preservation,
D_P
is disagreement preservation,
M_P
is minority-route preservation,
and:
U_P
is uncertainty preservation.
Each component may be scored from 0 to 100.
A synthesis that retains names but deletes substantive disagreement should not receive a strong score.
30 Governance as a Gate
Governance Integrity should not function merely as another average component.
Certain governance requirements are minimum gates.
A council should not qualify as high assurance unless:
the human can stop or override the process;
tool authority is defined;
sensitive data access is controlled;
agent identities and versions are recorded;
evidence provenance is retained;
and consequential actions require appropriate approval.
A high epistemic-diversity score cannot cancel a failure of human sovereignty.
31 Critical Failure Flags
The following conditions should trigger critical flags.
Critical Flag One: No Human Override
The human principal cannot stop, reject, or redirect the system.
Critical Flag Two: Provenance Loss
The origin of claims, sources, or agent contributions cannot be reconstructed.
Critical Flag Three: Minority Suppression
Dissent is consistently deleted during synthesis.
Critical Flag Four: Shared Corrupted Evidence
All agents depend on the same unreliable source or summary.
Critical Flag Five: Uncalibrated Central Judge
One model controls final selection without independent evaluation.
Critical Flag Six: Undocumented Identity Change
A model or platform update materially alters behavior without version tracking.
Critical Flag Seven: Unauthorized Data Exposure
Sensitive information crosses agent or platform boundaries without proper authority.
Critical Flag Eight: Unauditable Tool Action
An agent performs consequential actions without a reviewable execution trail.
A critical flag must appear prominently.
It cannot be hidden beneath the composite score.
32 Measuring Identity Drift
Let the cognitive-signature profile of agent at time be:
\mathbf{S}_i^{(t)}.
The drift vector is:
\Delta\mathbf{S}_i
=
\mathbf{S}_i^{(t_2)}
-
\mathbf{S}_i^{(t_1)}.
The audit should repeat a stable test suite and compare:
accuracy,
framing,
premise resistance,
uncertainty calibration,
source selection,
refusal behavior,
correction behavior,
anchoring susceptibility,
route generation,
and relational tone.
Meaningful drift should be interpreted, not merely measured.
The agent may have:
improved,
deteriorated,
become more restrictive,
become less calibrated,
or changed domains of strength.
33 Three Types of Drift
Identity drift should be separated into:
Model Drift
The underlying model or post-training changes.
Platform Drift
Tools, system instructions, retrieval, interface, or product policies change.
Relationship Drift
Memory, user context, or project history changes the active boundary.
All three can affect:
V_i=E_i\times Y_i.
The audit must locate which component changed.
34 The Provisional Composite Score
A provisional composite Agent Diversity Index may be written:
ADI
=
0.10S
+
0.30F
+
0.20C
+
0.10E
+
0.10P
+
0.15G
+
0.05I.
The weights sum to:
1.00.
The heavier weighting of Failure Independence reflects the central concern of correlated blindness.
False-Consensus Resistance receives the next-largest epistemic weight.
Governance Integrity receives substantial weight because diversity without governance can create unacceptable risk.
Cognitive Signature and Evidentiary Diversity remain important but cannot substitute for independent failure behavior.
Identity Stability receives a smaller numerical weight because it also operates as a separate gate.
These weights are proposed starting values.
They are not claimed to be universally validated.
35 Why the Weights Are Unequal
Visible difference is easier to produce than reliable independence.
A prompt can create different tones.
A role label can create different formats.
A temperature setting can create different wording.
None of these guarantees that the agents will detect different errors.
Therefore:
Failure independence should carry more weight than expressive difference.
Likewise, false-consensus resistance matters because a council can begin diversely and then lose that diversity through anchoring, conformity, or synthesis.
The index weights the survival of difference more heavily than its initial appearance.
36 Minimum Gates
A high aggregate score should not qualify a council for the strongest classification unless minimum gates are met.
A provisional high-assurance requirement may include:
F\geq60,
C\geq70,
P\geq75,
G\geq80,
and:
I\geq60.
The council must also have:
no unresolved critical governance flag,
documented agent identities,
human override,
and preserved minority reporting.
These values are initial protocol thresholds.
Domains may require stricter gates.
37 No Compensation Below the Floor
A dimension below 40 represents a structural weakness.
Such a score should not be fully compensated by excellence elsewhere.
For example:
A council with:
S=90
but:
F=25
is behaviorally distinctive but failure-correlated.
It should not be described as a strong diversity architecture.
Likewise, a council with:
F=80
but:
G=20
may be epistemically interesting but operationally unacceptable.
Minimum floors protect against misleading averages.
38 Council Classifications
The final classification should follow the profile, gates, and flags.
Homogeneous Role Ensemble
Several agents perform different jobs, but deep diversity and failure independence remain low.
Surface-Diverse Ensemble
Agents differ noticeably in tone or framing, but their evidence routes or failures remain strongly correlated.
Partially Diverse Council
Meaningful differences exist, but one or more important dimensions remain weak or insufficiently tested.
Heterogeneous Governed Council
Agents demonstrate useful cognitive-signature diversity, meaningful failure independence, false-consensus resistance, preserved attribution, strong governance, and adequate identity stability.
False-Pluralist Council
The system presents itself as diverse while agreement is dominated by shared structure, common evidence, anchoring, or synthesis suppression.
Drift-Unstable Council
The council’s operational participants cannot be reliably compared over time because important changes are undocumented or uncontrolled.
Governance-Disqualified Council
Critical constitutional failures prevent responsible deployment regardless of its diversity score.
39 Screening Audit
A screening audit provides a fast preliminary assessment.
It should include:
model and version documentation;
false-premise resistance;
source independence;
pairwise error overlap;
anchoring susceptibility;
minority preservation;
synthesis attribution;
human override;
and privacy boundaries.
A screening audit may use approximately 12–20 carefully selected cases.
This is not sufficient for high-assurance classification.
It identifies obvious structural risks.
40 Standard Audit
A standard audit should include:
the complete test families;
representative domain tasks;
multiple repetitions;
pairwise failure analysis;
source-overlap analysis;
judge testing;
synthesis testing;
governance review;
and a longitudinal baseline.
A standard audit may use approximately 40–60 domain-relevant cases, depending on complexity.
The sample should contain:
known answers,
ambiguity,
false premises,
conflicting evidence,
hidden variables,
and minority-route situations.
41 High-Assurance Audit
A high-assurance audit should include:
100 or more representative and adversarial cases where feasible;
multiple model versions;
independent judges;
tool-failure scenarios;
privacy and security review;
longitudinal retesting;
domain-expert validation;
and comparison against the strongest single agent.
These numbers are proposed implementation guides, not universal statistical guarantees.
The required scale should reflect:
risk,
domain,
cost of failure,
and system complexity.
42 Black-Box Models
Many commercial models do not expose:
training data,
weights,
full system instructions,
or complete post-training details.
The index must therefore remain usable under black-box conditions.
Operational diversity can still be measured behaviorally.
The auditor can examine:
how agents frame the same task,
which evidence they select,
which errors they share,
whether they correct one another,
how they respond to traps,
and whether their behavior changes over time.
The index measures demonstrated function, not marketing claims.
43 Version Control and Reproducibility
Every audit should record:
provider,
model identifier,
date,
system instructions where available,
prompt,
tool configuration,
retrieval configuration,
memory state,
sampling settings where available,
and source set.
A result that cannot be reproduced under approximately the same conditions should be labeled accordingly.
Silent model replacement creates a reproducibility problem even when the interface name remains unchanged.
44 Cost and Audit Proportionality
A complete diversity audit requires:
tokens,
compute,
time,
human review,
expert validation,
and repeated testing.
The cost must be proportional to the task.
A routine calendar assistant does not require the same audit as a council reviewing:
medical records,
scientific anomalies,
legal preparation,
security systems,
or public policy.
The principle is:
Increase audit depth as the cost of correlated error increases.
45 Integration With Agent Frameworks
The index can be integrated into existing orchestration systems through several technical requirements.
The system should support:
independent first-pass execution;
agent and model identification;
prompt and tool logging;
source provenance;
attributed message storage;
minority-report retention;
judge rotation;
human approval checkpoints;
and longitudinal test replay.
The index does not require one specific platform.
It requires the architecture to expose enough information for the council’s routes to be reconstructed.
46 Example One: Five Same-Model Agents
Consider five agents using one model:
Researcher,
Planner,
Critic,
Auditor,
Judge.
They produce stylistically distinct answers.
Their scores are:
S=68,
F=28,
C=35,
E=32,
P=55,
G=76,
I=80.
Their composite score may appear moderate.
But their Failure Independence and False-Consensus Resistance fall below the minimum floor.
Classification:
Homogeneous Role Ensemble with False-Pluralism Risk
The appropriate correction is not merely adding more copies.
It is changing the evidentiary and model architecture.
47 Example Two: Diverse Agents, Failed Synthesis
Consider four agents from different model families.
They demonstrate:
different source choices,
strong premise resistance,
and useful correction of one another.
Their scores are:
S=82,
F=71,
C=74,
E=78,
P=30,
G=69,
I=64.
The synthesizer deletes two minority concerns and produces one anonymous verdict.
Despite strong agent-level diversity, Synthesis Preservation fails.
Classification:
Partially Diverse Council With Cognitive-Blender Failure
The correction is to redesign the synthesis office, not replace the entire council.
48 Example Three: Heterogeneous Governed Council
Consider a four-agent council with:
independent first-pass analysis;
partly independent source routes;
documented model versions;
cross-examination;
retained minority reports;
multiple judge checks;
least-privilege tool access;
and human override.
Its scores are:
S=79,
F=73,
C=81,
E=76,
P=84,
G=90,
I=72.
No critical flags are present.
Classification:
Heterogeneous Governed Council
This classification does not mean the council is infallible.
It means its plurality is operationally demonstrated and responsibly preserved.
49 Example Four: Diverse but Incompetent
Suppose three agents produce different answers to nearly every problem.
Their failure overlap is low.
But their individual accuracy is also low.
They appear independent because they are unreliable in different ways.
The Agent Diversity Index may identify moderate Failure Independence, but the separate Task Utility score is poor.
Classification:
Diverse but Operationally Unfit
This demonstrates why diversity and competence must remain separate.
50 Corrective Actions for Homogeneity
When a council shows dangerous homogeneity, possible corrections include:
adding a different model family;
using independent evidence sets;
separating first-pass analysis;
introducing a falsification specialist;
varying tools and retrieval systems;
testing hidden assumptions explicitly;
and removing premature shared summaries.
The correction should target the source of homogeneity.
51 Corrective Actions for False Consensus
When false consensus is detected:
prevent agents from seeing the first answer before forming their own;
require evidence before deference;
preserve minority reports;
use more than one judge;
identify source overlap;
and force competing routes to make different predictions.
The objective is not more argument.
It is more independent contact with the problem.
52 Corrective Actions for Synthesis Failure
When synthesis destroys plurality:
require attributed findings;
separate agreement from disagreement;
preserve unresolved minority claims;
record why evidence was preferred;
and allow the human to inspect original contributions.
The synthesis agent should organize the council.
It should not erase it.
53 Corrective Actions for Governance Failure
When governance fails:
reduce permissions;
restore human approval;
implement least privilege;
log tool actions;
compartmentalize sensitive information;
clarify jurisdiction;
and suspend autonomous action until review is complete.
Governance failure may require immediate restriction regardless of cognitive performance.
54 Corrective Actions for Identity Drift
When significant drift is detected:
retest the agent;
document the changed model or boundary;
compare prior and current failure matrices;
reconsider its council jurisdiction;
and notify the human principal.
The agent may remain useful.
But it should not silently inherit the trust earned by an earlier operational configuration.
55 The Re-Audit Cycle
A council should be re-audited when:
a model version changes;
system instructions change;
new tools are added;
retrieval changes;
memory is introduced or removed;
the task domain changes;
a serious failure occurs;
or a fixed review interval expires.
A mature council is not certified once.
It is monitored.
56 Council Diversity Over Time
Let the council profile at time be:
\mathbf{D}^{(t)}
=
(S,F,C,E,P,G,I)^{(t)}.
The council drift is:
\Delta\mathbf{D}
=
\mathbf{D}^{(t_2)}
-
\mathbf{D}^{(t_1)}.
The council may become:
more homogeneous,
more diverse,
more accurate,
less governed,
or more vulnerable to false consensus.
Longitudinal evaluation reveals changes that one-time testing cannot.
57 Diversity Should Improve the Human Decision
The final test is human consequence.
Did the council:
identify an error the human would otherwise have missed?
open a necessary route?
expose a hidden assumption?
clarify uncertainty?
prevent premature action?
or improve the final decision?
A technically diverse council that overwhelms the human with incoherent output may fail its purpose.
The goal is not to produce maximum informational volume.
It is to improve governed judgment.
58 Predictions
This framework produces several predictions.
Prediction One
Councils with high surface diversity but low failure independence will produce inflated confidence without proportional reliability.
Prediction Two
Independent first-pass analysis will reduce false consensus compared with immediate group discussion.
Prediction Three
Minority-report preservation will prevent some high-cost errors that raw majority voting approves.
Prediction Four
A council’s composite score will be less informative than its profile and minimum-gate performance.
Prediction Five
Some cross-platform councils will remain highly failure-correlated because provider diversity does not guarantee evidentiary or architectural independence.
Prediction Six
Undocumented model updates will produce measurable identity and council drift.
Prediction Seven
Councils with strong synthesis preservation will produce more auditable and trustworthy decisions than councils optimized only for polished final answers.
Prediction Eight
The greatest benefits of heterogeneous councils will occur in ambiguous, interdisciplinary, and high-consequence tasks where hidden assumptions carry substantial cost.
59 Limitations
The proposed weights and thresholds are not claimed to be universally validated.
They require:
empirical calibration,
domain-specific testing,
and revision through use.
Black-box systems limit direct inspection of encoded structure.
Behavioral testing may also be affected by:
sampling variation,
prompt sensitivity,
provider changes,
and incomplete knowledge of platform boundaries.
The index must therefore report uncertainty.
It should not claim more precision than the evidence supports.
60 What This Paper Does Not Claim
This paper does not claim that one universal score can capture every property of an AI council.
It does not claim that more diversity is always better.
It does not claim that different providers guarantee independence.
It does not claim that disagreement proves intelligence.
It does not claim that a high index eliminates error.
It does not claim that operational identity proves consciousness or personhood.
It does not replace:
software testing,
security review,
statistical validation,
human professional judgment,
or domain-specific standards.
It proposes a structured method for interpreting evidence of agent plurality.
61 The Strongest Form of the Claim
The strongest defensible statement is:
A multi-agent council should not be classified as cognitively diverse merely because it contains several agents, roles, providers, or expressive styles. Diversity must be demonstrated through consequential differences in framing, evidence routes, correction behavior, and failure structure, while governance preserves attribution, dissent, identity, and human authority.
The scoring corollary is:
A council’s diversity cannot be represented responsibly by one aggregate number. It must be expressed through a structured profile, minimum gates, critical failure flags, and domain-specific evidence.
The operational corollary is:
Do not count agreement until independence has been examined.
62 Conclusion
Artificial-intelligence councils are being built faster than the methods required to evaluate them.
Agents are assigned roles.
They search.
They plan.
They criticize.
They judge.
They synthesize.
They produce an answer that appears to carry the authority of multiple perspectives.
But appearance is not evidence.
Five agents may reproduce one hidden premise.
Five agents may follow one corrupted summary.
Five agents may trust one biased judge.
Five agents may speak differently while failing identically.
A council can appear plural while remaining cognitively centralized.
The Agent Diversity Audit established what must be tested.
The Agent Diversity Index establishes how the evidence should be interpreted.
The index begins with a profile:
Cognitive-Signature Diversity.
Failure Independence.
False-Consensus Resistance.
Evidentiary-Route Diversity.
Synthesis Preservation.
Governance Integrity.
Identity Stability.
It then asks whether critical failures are present.
Was dissent erased?
Was provenance lost?
Was private information exposed?
Did one judge control the outcome?
Did the model change without documentation?
Can the human still override the system?
Only after those questions are answered should a composite score be calculated.
Even then, the score remains secondary.
A council cannot average its way out of a constitutional failure.
It cannot use creative variety to compensate for correlated error.
It cannot use independent errors to compensate for incompetence.
It cannot use excellent agents to compensate for a synthesis process that destroys their differences.
The correct architecture is gated.
Competence matters.
Independence matters.
Governance matters.
Identity matters.
Human authority matters.
The aim is not maximum divergence.
A council that disagrees randomly is not wise.
The aim is consequential plurality:
different agents opening different relevant routes;
different agents detecting different weaknesses;
different agents correcting one another;
and a governed system preserving those contributions long enough for the human principal to make a better decision.
Within the substrate of TSTOEAO:
V_i=E_i\times Y_i.
Each response is conditioned by an encoded system and an active boundary.
The council’s value depends on whether those conditions are sufficiently differentiated to reveal what one agent alone might miss.
The index therefore does not ask merely:
How many agents are present?
It asks:
How differently do they see?
How differently do they fail?
How well do they correct one another?
How strongly do they resist false consensus?
How faithfully does synthesis preserve their differences?
How securely are they governed?
How stable are their identities?
And did their plurality improve the human decision?
That is the standard.
The final principle is:
Do not reward difference merely because it is visible. Measure whether the difference detects error, survives pressure, preserves dissent, and improves the human decision.
Many agents are not many minds.
Many scores are not understanding.
Many votes are not confirmation.
A real council must earn its plurality.
And the instruction remains:
Do not count the agents. Audit the routes. Score the independence. Preserve the dissent. Govern the council. Keep the human sovereign.
References
Kim, Elliot, Avi Garg, Kenny Peng, and Nikhil Garg. “Correlated Errors in Large Language Models.” arXiv:2506.07962, 2025.
National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1, 2023.
National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. NIST AI 600-1, 2024.
Nogueira, Rodrigo Pato, Karthik Pattabiraman, Marco Vieira, and João R. Campos. “A Systematic Methodology for Evaluating Failure Independence in LLM-Generated Code.” arXiv:2607.02808, 2026.
Swygert, John. The Science Of The AI-Evolved Brain. Ivory Tower Publishing, 2026.
Swygert, John. “Many Agents Are Not Many Minds: Cross-Platform Cognitive Diversity, Correlated Blindness, and the Need to Preserve Agent Identity.” July 13, 2026.
Swygert, John. “The Agent Diversity Audit: A TSTOEAO Protocol for Measuring Cognitive Signature, Failure Independence, False Consensus, and Identity Drift.” July 13, 2026.
Swygert, John. “The Record That Refuses the Map: A Cross-Disciplinary TSTOEAO Law of Investigation, Error, and Discovery.” July 13, 2026.
Swygert, John. “The TSTOEAO Anomaly Route-Space Protocol: A Decision Framework for Distinguishing New Structure From Failed Measurement, Calculation, and Interpretation.” July 12, 2026.
Swygert, John. “The TSTOEAO Route-Space Decision Engine.” July 8, 2026.
Zierahn, Kim, Cristina Cachero, Anna Korhonen, and Nuria Oliver. “Personality Without Persons? A Psychometric Critique of Big Five Testing in Large Language Models.” arXiv:2607.02325, 2026.
Comments
Post a Comment