Skip to main content
Assets
iPulse AI Concepts
Consensus & Disagreement

I Audited 4,523 AI Judgments. Here’s What I’d Check Before Trusting an Agent Team.

Across 377 assets, a simple vote and a risk-aware scoring system disagreed in 45% of decisive cases. An audit for builders of agent teams.

Russlan RamdowarFounder of iPulse AI12 min read
Architecture diagram contrasting winner-only voting with preserved agent estimates, evidence, uncertainty measures, a versioned aggregation policy, abstention, and outcome evaluation. Audit callout: 166 of 365 decisive comparisons differed, rounded to 45%.
Multi-agent systemsAI consensusDecision evidenceAI uncertainty

Originally published in Level Up Coding on September 8, 2026. Republished on iPulse AI on September 19, 2026. Figures and findings retain the original research dates; this is a historical research edition.

Twelve AI agents. Twelve STRONG BUY ratings. What would you put on the screen?

Probably STRONG BUY. Perhaps a reassuring little badge: 100% agreement.

Now the awkward part. Three assets in our research snapshot had that identical ballot. Two finished STRONG BUY. The third finished BUY.

Same votes. Different verdict. So what was the vote not telling us?

Accenture and Adyen kept the stronger rating. The Trade Desk did not. Its agents hadn’t secretly changed their minds. Our scoring layer used information their labels could not carry: forecast magnitude, consistency, and risk.

That is the part I want anyone building an agent team to notice. You can collect every answer, count every vote correctly, and still throw away information that changes the decision.

The counting can be flawless. The summary can still be inadequate.

Think of twelve engineers saying a migration is ‘safe.’ One means a brief outage. Another means no outage but a painful rollback. A green badge makes them look unanimous. Ask what they mean by safe, and suddenly you have a meeting worth having.

I wanted to put a number on that gap.

We audited 4,523 agent judgments across 377 assets. In the 365 cases with a clear winning direction, a simple vote and the richer scoring system disagreed 166 times: about 45%.

No, that does not mean 45% of the forecasts were wrong. We compared two ways of summarizing model outputs, not predictions against future market results. But it does mean the choice of aggregator was consequential.

Before you pay for another agent, inspect what happens to the answers you already have. Here is the audit I would start with.

In an earlier investor-facing analysis, I examined how often these forecasts disagreed. Here, I reuse that snapshot to test a different engineering question: can the votes reproduce the final decision?

The snapshot comes from iPulse AI, the investment-research platform I build. Its live rankings can change; this audit uses the fixed July snapshot described below.

What you can steal from this audit

  • A baseline vote you can reproduce without another model call.
  • A transition matrix showing exactly where your aggregator changes the answer.
  • An identical-ballot test that exposes information your labels cannot preserve.
  • A way to separate agreement, disagreement, and actual confidence.
  • A decision record that lets you investigate a result instead of arguing with a screenshot.

First, count something embarrassingly simple

Our sample covered stocks, cryptoassets, commodities, indices, and currencies at a five-year forecast horizon. There were 376 assets with twelve analytical voices and one with eleven: 4,523 judgments in total. Each voice returned one of five ordered ratings, from strongly positive to strongly negative:

STRONG_BUY, BUY, NEUTRAL, PARTIALLY_SELL, SELL_ALL

For the baseline, I grouped STRONG_BUY and BUY as positive, kept NEUTRAL separate, and grouped PARTIALLY_SELL and SELL_ALL as negative. Whichever group had the most votes won:

positive = strong_buy + buy
negative = partially_sell + sell_all
if positive > max(neutral, negative):
    vote_direction = "POSITIVE"
elif negative > max(neutral, positive):
    vote_direction = "NEGATIVE"
elif neutral > max(positive, negative):
    vote_direction = "NEUTRAL"
else:
    vote_direction = "TIE"

Technically, this is a plurality, not necessarily a majority. Five positive, four neutral, and three negative votes produce a positive winner, although seven agents did not vote positive. A tied maximum stays a tie. That distinction already makes a ‘consensus’ badge less comforting.

The production aggregator had a different job. It retained continuous forecasts, volatility, directional consistency, dividend effects, and configured thresholds. I mapped its final ratings into the same three directions so the two methods could be compared.

The question was deliberately narrow: could the ballot reproduce the richer decision?

The answer was no. The interesting part was where it failed.

166 disagreements. Almost none were reversals.

Of 377 assets, twelve had a tied baseline vote. That left 365 decisive comparisons. In 166, the final direction differed. Here is what that 45% actually contained:

Matrix comparing naive positive, negative, neutral, and tied votes with final positive, neutral, and negative consensus directions. Ninety-eight positive pluralities and sixty-seven negative pluralities became neutral.

Figure 1. A naive plurality frequently became neutral after the scoring pipeline evaluated magnitude, consistency, volatility, and thresholds. Counts describe model outputs, not investment outcomes. Source: author-controlled production research snapshot, July five-year original snapshot.

Source: iPulse AI

Read the destinations, not just the headline:

  • 98 positive pluralities became neutral.
  • 67 negative pluralities became neutral.
  • One negative plurality became positive.
  • No positive plurality became negative.

So 165 of the 166 disagreements ended at neutral. This was mostly a directional vote meeting a scoring threshold it did not clear, not a system shouting BUY while another shouted SELL.

Important distinction: NEUTRAL is a completed financial rating. ABSTAIN means the system declines to produce a decision. These records show the former; they do not prove the latter. Nor does landing at neutral prove the threshold was correct.

Still, imagine replacing that scoring layer with ‘most votes wins.’ You would be changing the decision policy while calling it a simpler implementation. That’s quite a refactor.

Then I tried saving every vote. Still not enough.

Maybe I had created the problem by squeezing five ratings into three directions. Fair objection. So I checked the full five-category ballot.

It still could not reconstruct the final label.

The 377 assets contained 267 distinct vote patterns. Twenty-four patterns appeared alongside more than one final rating. Those patterns covered 77 assets, about 20% of the sample.

The unanimous cases from the opening make this easy to see:

Comparison table for Accenture, Adyen, and The Trade Desk. All received twelve STRONG BUY ratings, but The Trade Desk finished with a BUY consensus while the other two finished STRONG BUY.

Figure 2. Three assets had the same 12–0–0–0–0 histogram, but continuous scoring inputs produced different scores and two final labels. Values are model outputs from one governed snapshot, not recommendations.

Source: iPulse AI

The Trade Desk had the highest annualized mean forecast of these three and the strongest direction-consistency reading. It also had the highest risk pressure, the lowest continuous consensus score, and the weaker final label. Those are model outputs, not realized returns or investment recommendations.

Think of star ratings. Two hotels can both score four stars while differing wildly on price, noise, and accessibility. Counting twelve four-star reviews does not recover the details that matter to your trip. A forecast label makes the same kind of compression.

Counting the labels compresses the evidence a second time.

My practical test: group records by the complete vote pattern, then count distinct final verdicts within each group. If a group has several verdicts, your ballot is missing inputs needed to explain them. That can be intentional and sensible. It must not be invisible.

Please stop turning agreement into a probability

Eight of twelve agents choose BUY. Agreement: 67%. Probability of being right: we have not established one. Changing the field name to ‘confidence’ does not finish the experiment.

Now compare eight BUY plus four STRONG BUY with eight BUY plus four SELL ALL. Same winning share. Very different opposition. One group argues about enthusiasm; the other argues about direction.

We measured categorical spread using normalized Shannon entropy. It describes how votes distribute across the five categories. Crucially, it cannot distinguish those two eight-to-four ballots either: the category proportions are identical. You still need the labels and their meaning. For proportions p_i across K categories:

H_norm = -sum(p_i * ln(p_i)) / ln(K)
from math import log
def normalized_vote_entropy(counts):
    if len(counts) != 5:
        raise ValueError('Expected all five rating counts')
    if any(type(c) is not int or c < 0 for c in counts):
        raise ValueError('Counts must be nonnegative integers')
    total = sum(counts)
    if total == 0:
        raise ValueError('No votes; do not invent agreement')
    probabilities = [c / total for c in counts if c]
    return -sum(p * log(p) for p in probabilities) / log(5)

Unanimity gives zero entropy. A more evenly spread ballot gives a higher value. Neither tells you whether the answer is correct. The denominator stays at five even when some categories have no votes.

Across our sample, however, higher entropy travelled with a noticeably different decision state:

Four bar charts showing that higher vote-entropy quartiles have more neutral final signals, wider return dispersion, lower direction consistency, and higher risk pressure.

Figure 3. Higher five-class vote entropy coincided with more neutral final signals, wider forecast dispersion, lower directional consistency, and higher risk pressure. Quartiles contain 94 or 95 assets. Associations are descriptive, not causal.

Source: iPulse AI

In the lowest-entropy quarter, 5% of assets finished neutral. In the highest, 83% did. Average annual-return dispersion widened from roughly four percentage points to eight. Direction consistency fell from 0.84 to 0.71; risk pressure rose from 40 to 58.

These are associations inside one scoring system, not proof that entropy caused neutrality. My takeaway is simpler: keep uncertainty visible. One cheerful percentage is doing too many jobs.

What I’d change before adding agent thirteen

I would start with the handoff between agents and the application. The architecture diagram above shows the proposed flow. These five changes make that handoff inspectable; they are design recommendations, not claims that every feature was tested in this audit.

1. Keep the receipt, not just the verdict

Store each agent’s estimate, units, horizon, evidence timestamp, model version, and source references. This illustrative record is a starting point, not a production row:

{
  "agent_id": "value_v3",
  "rating": "BUY",
  "estimate": 0.14,
  "estimate_unit": "annualized_total_return_fraction",
  "horizon": "5Y",
  "evidence_as_of": "<timestamp>",
  "model_version": "<version>",
  "source_ids": ["<source-reference>"]
}

Otherwise, six months later you will know that an agent said BUY, but not what it was buying into. A string is cheap to store. An explanation is expensive to reconstruct.

2. Separate the questions your confidence badge hides

Keep distinct fields for:

  • Vote concentration: how tightly do the categories cluster?
  • Forecast dispersion: how far apart are the numerical estimates?
  • Directional consistency: do the modeled paths point the same way?
  • Dependence: how much model or source material is shared?
  • Freshness: what evidence is stale, absent, or incompatible?
  • Action cost: what happens if this decision is wrong?

Twelve personas using the same model and sources are not automatically twelve independent experts. Agent count tells you how many outputs arrived, not how many independent checks occurred. Separately, Google Research’s study of 180 agent configurations found that performance depended on architecture and task structure. Headcount alone is a poor design brief.

3. Make the decision policy something you can replay

Where comparable numerical estimates exist, test aggregating them before applying category thresholds. Keep risk rules explicit. Do not average incompatible horizons or units and hope that a persuasive summary will rescue the result.

Version the thresholds alongside the inputs. Otherwise, ‘the model changed its mind’ might really mean ‘we moved a boundary.’ Those deserve different investigations.

4. Give ‘not enough evidence’ its own exit

A system should be able to ask for missing evidence, route a case to a human, or abstain. A neutral rating cannot substitute for that branch: ‘I evaluated this and got neutral’ differs from ‘I could not evaluate this reliably.’

Specify the exit conditions before the result reaches an action button. Selective-classification research provides the useful framing: evaluate both the error on answered cases and how many cases the system answers.

5. Return a decision record you can investigate

For example, an illustrative abstention record could separate the process outcome from the domain rating:

{
  "decision_id": "<stable-id>",
  "decision_type": "ABSTAIN",
  "rating": null,
  "reasons": ["MISSING_REQUIRED_EVIDENCE"],
  "agent_records_ref": "<immutable-record-set>",
  "aggregation_version": "<version>",
  "decided_at": "<timestamp>"
}

The referenced record set should retain votes, continuous estimates, disagreement measures, and provenance. When the rating is completed, use a different decision type and populate it. Do not send both a normal rating and an unexplained abstention flag to a consumer and let it guess.

Fine. But does any of this make the answers better?

That is the question an enthusiastic architecture diagram cannot answer. More fields make investigation possible. They do not automatically make a forecast more accurate.

The next test has to use outcomes unavailable at decision time:

  1. Freeze agent outputs, evidence timestamps, and model and policy versions.
  2. Replay simple voting and richer aggregation on the same eligible cases.
  3. Measure directional and numerical forecast errors separately.
  4. For abstaining variants, report error alongside coverage: the proportion of cases answered.
  5. Check sensitivity to prompt, retrieval, model, and threshold changes.
  6. Include latency and cost. A better result can still be unusable if it arrives too late.

A system that answers one easy question and refuses the next ninety-nine can advertise beautiful accuracy. You probably would not hire it. That is why coverage belongs beside the score, not buried in a footnote.

Run the candidate in shadow mode first: record its proposed decisions without letting them trigger actions. Compare later against matured outcomes. Keep the original evidence frozen so hindsight cannot quietly rewrite the test.

I would rather see an honest comparison lose than a sophisticated aggregator win a contest designed around itself.

Open one trace before your next agent meeting

You can start without retraining anything. Pick one completed decision and ask:

  1. Can I recover every input, or only the winner?
  2. Does ‘confidence’ mean measured correctness, or just a vote fraction?
  3. Would identical ballots with different magnitudes or risks be distinguishable here?
  4. Can I tell a completed neutral decision from a refusal to decide?

One trace will not validate the whole system. It can reveal that you lack the records required to validate it at all. That is a useful discovery to make before the incident review.

The boundary around our result

This audit covers one five-year financial-research snapshot. The separately configured voices are not statistically independent samples. Dispersion and risk measures are model-derived, and the scoring policy itself helps determine the patterns observed.

The 45% is disagreement between methods across 365 decisive cases. It is not a failure rate across 4,523 judgments, a trading return, or evidence that either method outperformed the market.

What the audit does establish is that the ballot did not contain enough information to reproduce the richer decision. Whether those extra distinctions improve future outcomes remains an empirical question.

Back to those twelve green lights

Accenture. Adyen. The Trade Desk. Twelve STRONG BUY votes each. Two final rating levels. The puzzle disappears once you inspect what the labels left out.

That is the lesson I would carry into an agent-based code reviewer, approval workflow, or research assistant. Before celebrating agreement, ask what the application will never see.

The uncomfortable engineering question is not always ‘Which model should we add?’ Sometimes it is ‘Why did we pay for all this analysis and then keep only the label?’

Count the votes. Keep the evidence.

Then earn the right to act.

Which field would you most hate to discover missing from your agent logs after a bad decision?

Resources

Research note

Source: an author-controlled July 2026 research snapshot covering 377 assets and 4,523 agent judgments at a five-year horizon. The analysis keeps the latest active version of each eligible asset/configuration, compares three-way vote winners with final model directions, and measures entropy across all five rating categories. Twelve tied votes are excluded from the 365-case disagreement denominator. Percentages are rounded. Reproduction SQL and analysis definitions are retained with the research package. This is an engineering case study, not personalized investment advice.

Evidence register

Empirical sources behind this publication

Concept publication record

This owned edition preserves the stable iPulse AI concept record, including visible corrections and source lineage. It is editorial analysis, not formal research, a guarantee, or personalized investment advice.

View the Level Up Coding edition

Follow iPulse AI concepts

Receive the next iPulse AI investigation by email.

Subscribe on Substack

Continue the research

Related iPulse AI publications