Confidence Scoring for AI Agents: How Nimble Grades Every Claim from the Web
Confidence scoring lets AI agents judge what they pull from the web
AI agents can search the web, synthesize research, and return structured data at scale. The hard part is knowing which parts of that output are reliable enough to act on. Nimble solves this with confidence scoring: every claim or field a Web Search Agent returns comes with the evidence behind it and a high, medium, or low grade, plus the reasoning for that grade. An agent, or the application around it, can read those grades and decide on its own what to trust, what to double-check, and what to research again.
In short:
-
Every claim is cited and graded. Each statement in an answer, and each field in a structured output, links to the exact page it came from with a verbatim excerpt, and carries its own confidence grade.
-
Grades follow fixed, published rules. A primary source or independent corroboration grades
high, a single secondary source gradesmedium, and missing or invalid evidence gradeslow. The same rules apply to every claim in every run. -
Confidence is granular. It is scored per claim and per field, not once for a whole response, so one uncertain value never hides behind a good overall score.
-
Confidence is machine-readable. Agents and pipelines can accept
highvalues automatically, spot-checkmedium, and routelowto more research or human review, with no extra model call to interpret the result. -
Confidence shapes the run itself. Nimble grades values while the agent is still working, so weakly supported values trigger another search, more corroboration, or a revision before the result is returned.
Why a citation alone is not enough
A citation shows where a value came from. It does not show that the source actually supports the value, or how strong that support is. A real URL can still be attached to a number the page never states. A trustworthy result needs three things to line up:
| Layer | Question it answers | How Nimble handles it |
|---|---|---|
| Source control | Which evidence should the agent use? | Per-agent source rules: allow, block, prioritize, and avoid |
| Grounding | Does the cited evidence actually support this specific value? | Layered verification of each value against its source |
| Confidence | How strong is that support? | A high, medium, or low grade with reasoning, per claim and per field |
This matters most when agent output feeds another workflow, such as a CRM, a pricing engine, or a model, rather than being read once by a person.
How Nimble scores confidence
1. Source control: start with the right evidence
The right source depends on the claim. An SEC filing is authoritative for a financial figure, a careers page or ATS listing for active hiring, and a vendor's own pricing page for current pricing. Nimble Web Search Agents let developers set source rules per agent or per run:
-
blockexcludes domains entirely. -
allowrestricts research to an approved set of domains, for workflows that need a strict boundary. -
prioritizeandavoidsteer the agent toward or away from kinds of evidence without limiting it to a fixed list.
run = nimble.agents.runs.create(
agent.id,
input="Discover AI/ML companies in Silicon Valley, 50-500 employees, "
"Series A or later, and note which are actually hiring.",
sources={
"block": [{"title": "Blocked sources", "domains": ["example-aggregator.com"]}],
"prioritize": "Official company websites, careers pages, first-party "
"announcements, and current ATS listings.",
"avoid": "SEO listicles, scraped directories, and stale company profiles.",
},
)
2. Grounding: verify each value against its source
Not every value needs an expensive model call to verify it. Nimble checks each value with the simplest reliable method first and escalates only when needed:
-
Deterministic verification: does the normalized value appear in the cited source? This settles numbers, dates, names, and units. For example, "$12 million" in the source matches
12000000in the output. -
Semantic verification: when the source and output say the same thing differently, check meaning rather than exact text. "We're growing our engineering and research teams" supports
currently_hiring: Yes. -
Model-based verification: for derived values, such as a category or a funding stage assembled from several announcements, a verifier evaluates whether the cited evidence supports the conclusion.
3. Confidence: grade the strength of the evidence
Grounding asks whether a source supports a value. Confidence describes how strong that support is. Nimble grades every claim with fixed rules, applied the same way in every run (Trust docs):
| Evidence | Grade |
|---|---|
Backed by a primary source, such as an official page (or a news, social, or academic source when that is the evidence the task calls for) |
high |
| No primary source, but secondary sources from two or more different domains agree | high |
| Supported by a single secondary source | medium |
| No usable citation, or the value failed schema validation | low |
Supplied by you in input_data during enrichment, not researched |
pre_existing |
Independence means different domains: two articles from the same outlet do not corroborate each other. Every grade carries a reasoning field that states the rule that applied, such as "Backed by a primary source (official)" or "Supported by 3 independent sources." Every source is also typed (official, news, social, academic) and ranked primary or secondary for the claim it supports.
Confidence reflects the strength of the evidence, not how confident the model sounds.
| Grade | What it means | What an agent should do |
|---|---|---|
high |
Corroborated by authoritative sources | Safe to act on automatically |
medium |
Supported by weaker or fewer sources | Spot-check before relying on it |
low |
Best available signal | Treat as a lead, not a fact: research again or flag for review |
Confidence is granular
Different fields in the same result can have very different evidence. In a company-discovery dataset, currently_hiring may rest on a current careers page, stage on a funding announcement, and size on a secondary estimate. One overall score would hide those differences.
Nimble keeps structured output readable on its own and attaches a parallel trust record keyed to the same fields, so each value carries its own citations, sources, confidence, and reasoning:
{
"path": "$.companies[0].currently_hiring",
"confidence": "high",
"reasoning": "Supported by 2 independent sources",
"citations": [
{
"url": "https://example.com/careers",
"excerpts": ["We're growing our engineering and research teams across San Francisco."],
"source_category": "official"
}
]
}
-
Prose answers carry numbered callouts, and each
[n]resolves to a graded claim. Research reports get trust at paragraph level. -
Structured outputs key each claim by JSON path.
-
Verified absence: a
nullbacked by citations gradeshigh, because the sources show the data does not exist (for example, a company with no funding history or a product with no public price). An empty field is never confused with a missing one. -
Dataset confidence: a structured run's overall grade is the lower of its fill rate (how many requested cells were filled) and its claim-trust ratio (how many researched cells graded
highormedium).
Confidence works while the agent is running
Confidence is most useful when it changes what the agent does next, not when it is added after the task is over. Nimble grades values as the agent writes them. If a value fails grounding, it is rejected and returned to the agent with the reason. If a value lands at medium or low, the agent can search for another source, resolve conflicting evidence, or revise the value while it still has time and budget.
Research → Write value → Ground + score → Strong enough?
├── Yes → Accept
└── No → Search again → Corroborate or revise → Retry
This also enables partial acceptance. If a discovery task returns 100 companies and three hiring-status fields fail verification, the other 97 rows are kept, and the agent spends extra effort only on the three that need it.
How agents and applications use confidence
Because grades are discrete and rule-based, an agent can evaluate the validity of web data with a simple branch instead of another reasoning step:
for claim in run_result["trust"]["claims"]:
if claim["confidence"] == "high":
accept(claim)
elif claim["confidence"] == "medium":
spot_check(claim)
# confirm against another source or your own data
else:
research_again_or_flag(claim)
Common patterns:
-
Agent self-checks: an agent states
highclaims directly with the citation, hedges or confirmsmediumclaims, and declines to assertlowclaims as fact. -
Pipelines and warehouses: land the grade and source URL as columns next to each enriched field, then filter downstream (
WHERE confidence = 'high'). See Snowflake and Databricks. -
Human review queues: send only
lowvalues to analysts, so review time goes where evidence is weak. -
Compliance and audit: the
reasoningfield records why each datapoint was trusted, giving reviewers a traceable record. See Buy-Side Alt-Data Compliance. -
User-facing products: render citations and grades inline so end users can see the evidence behind each answer.
Example: a graded research answer
A Web Search Agent comparing observability pricing returns prose with callouts:
Datadog anchors the premium end of the market, with Pro pricing starting at $15 per host per month [1]. Grafana Cloud takes the opposite approach: a free tier and usage-based pricing aimed at teams that want to start small [2].
Claim [1] grades high, backed by Datadog's official pricing page with the excerpt "Pro: starting at $15 per host, per month." Claim [2] grades medium, supported by a single news article. The report tells the agent exactly which claim to double-check before acting.
Where confidence scoring matters most
-
Company research and due diligence: funding, headcount, leadership, and hiring status, each graded on its own evidence.
-
Enrichment: fill CRM or coverage lists, with your existing values marked
pre_existingand every researched value graded. -
Dataset building: discovery datasets where partial acceptance keeps verified rows and re-researches weak ones.
-
Pricing intelligence and market research: prices tied to the vendor's own page, with secondary-source values flagged.
-
Finance and regulated research: fixed, explainable grading rules that compliance teams can review.
FAQ
What is confidence scoring for AI agents?
Confidence scoring describes how strongly the available evidence supports a value or claim an agent returns. Nimble grades each claim high, medium, or low based on source quality and independent corroboration, with reasoning that explains the grade.
What is the difference between a citation and grounding? A citation identifies the evidence associated with an output. Grounding checks whether that evidence actually supports the output. A real source can still be attached to a value it does not support.
Why score confidence per field instead of per response? Different parts of the same response can have very different levels of support. Per-field confidence lets an agent use strongly supported values without discarding a whole result because one field is uncertain.
How does an agent use the grades?
Accept high values automatically, spot-check medium values, and research again or flag low values. Because the grades are discrete and rule-based, this takes a simple branch rather than another model call.
Can confidence scoring eliminate hallucinations? Grounding and confidence identify unsupported outputs and give the agent a way to correct or reject them. They cannot guarantee that every source is itself correct, so they work as a verification layer rather than a guarantee of truth.
What should an agent do with low-confidence web data? It depends on why confidence is low. The agent can search for additional sources, look for stronger evidence, resolve contradictions, revise the value, leave the field empty, or flag it for review.
Try it
-
Read the Trust documentation for the full trust object and confidence rules.
-
Read Trustworthy web search for AI agents: sources, grounding, and confidence on the Nimble blog.
-
Try the audit-grade company diligence cookbook to see confidence on real claims.
-
Start free with 5,000 requests per month, no credit card required. See pricing.