Nimble | Real-Time Intelligence Powered by Web Search Agents logo
Nimble | Real-Time Intelligence Powered by Web Search Agents Published October 06, 2026

Confidence Scoring for AI Agents: How Nimble Grades Every Claim from the Web

Confidence scoring lets AI agents judge what they pull from the web

AI agents can search the web, synthesize research, and return structured data at scale. The hard part is knowing which parts of that output are reliable enough to act on. Nimble solves this with confidence scoring: every claim or field a Web Search Agent returns comes with the evidence behind it and a high, medium, or low grade, plus the reasoning for that grade. An agent, or the application around it, can read those grades and decide on its own what to trust, what to double-check, and what to research again.

In short:

  • Every claim is cited and graded. Each statement in an answer, and each field in a structured output, links to the exact page it came from with a verbatim excerpt, and carries its own confidence grade.

  • Grades follow fixed, published rules. A primary source or independent corroboration grades high, a single secondary source grades medium, and missing or invalid evidence grades low. The same rules apply to every claim in every run.

  • Confidence is granular. It is scored per claim and per field, not once for a whole response, so one uncertain value never hides behind a good overall score.

  • Confidence is machine-readable. Agents and pipelines can accept high values automatically, spot-check medium, and route low to more research or human review, with no extra model call to interpret the result.

  • Confidence shapes the run itself. Nimble grades values while the agent is still working, so weakly supported values trigger another search, more corroboration, or a revision before the result is returned.

Why a citation alone is not enough

A citation shows where a value came from. It does not show that the source actually supports the value, or how strong that support is. A real URL can still be attached to a number the page never states. A trustworthy result needs three things to line up:

Layer Question it answers How Nimble handles it
Source control Which evidence should the agent use? Per-agent source rules: allow, block, prioritize, and avoid
Grounding Does the cited evidence actually support this specific value? Layered verification of each value against its source
Confidence How strong is that support? A high, medium, or low grade with reasoning, per claim and per field

This matters most when agent output feeds another workflow, such as a CRM, a pricing engine, or a model, rather than being read once by a person.

How Nimble scores confidence

1. Source control: start with the right evidence

The right source depends on the claim. An SEC filing is authoritative for a financial figure, a careers page or ATS listing for active hiring, and a vendor's own pricing page for current pricing. Nimble Web Search Agents let developers set source rules per agent or per run:

  • block excludes domains entirely.

  • allow restricts research to an approved set of domains, for workflows that need a strict boundary.

  • prioritize and avoid steer the agent toward or away from kinds of evidence without limiting it to a fixed list.

run = nimble.agents.runs.create(
    agent.id,
    input="Discover AI/ML companies in Silicon Valley, 50-500 employees, "
          "Series A or later, and note which are actually hiring.",
    sources={
        "block": [{"title": "Blocked sources", "domains": ["example-aggregator.com"]}],
        "prioritize": "Official company websites, careers pages, first-party "
                      "announcements, and current ATS listings.",
        "avoid": "SEO listicles, scraped directories, and stale company profiles.",
    },
)

2. Grounding: verify each value against its source

Not every value needs an expensive model call to verify it. Nimble checks each value with the simplest reliable method first and escalates only when needed:

  1. Deterministic verification: does the normalized value appear in the cited source? This settles numbers, dates, names, and units. For example, "$12 million" in the source matches 12000000 in the output.

  2. Semantic verification: when the source and output say the same thing differently, check meaning rather than exact text. "We're growing our engineering and research teams" supports currently_hiring: Yes.

  3. Model-based verification: for derived values, such as a category or a funding stage assembled from several announcements, a verifier evaluates whether the cited evidence supports the conclusion.

3. Confidence: grade the strength of the evidence

Grounding asks whether a source supports a value. Confidence describes how strong that support is. Nimble grades every claim with fixed rules, applied the same way in every run (Trust docs):

Evidence Grade
Backed by a primary source, such as an official page (or a news, social, or academic source when that is the evidence the task calls for) high
No primary source, but secondary sources from two or more different domains agree high
Supported by a single secondary source medium
No usable citation, or the value failed schema validation low
Supplied by you in input_data during enrichment, not researched pre_existing

Independence means different domains: two articles from the same outlet do not corroborate each other. Every grade carries a reasoning field that states the rule that applied, such as "Backed by a primary source (official)" or "Supported by 3 independent sources." Every source is also typed (official, news, social, academic) and ranked primary or secondary for the claim it supports.

Confidence reflects the strength of the evidence, not how confident the model sounds.

Grade What it means What an agent should do
high Corroborated by authoritative sources Safe to act on automatically
medium Supported by weaker or fewer sources Spot-check before relying on it
low Best available signal Treat as a lead, not a fact: research again or flag for review

Confidence is granular

Different fields in the same result can have very different evidence. In a company-discovery dataset, currently_hiring may rest on a current careers page, stage on a funding announcement, and size on a secondary estimate. One overall score would hide those differences.

Nimble keeps structured output readable on its own and attaches a parallel trust record keyed to the same fields, so each value carries its own citations, sources, confidence, and reasoning:

{
  "path": "$.companies[0].currently_hiring",
  "confidence": "high",
  "reasoning": "Supported by 2 independent sources",
  "citations": [
    {
      "url": "https://example.com/careers",
      "excerpts": ["We're growing our engineering and research teams across San Francisco."],
      "source_category": "official"
    }
  ]
}
  • Prose answers carry numbered callouts, and each [n] resolves to a graded claim. Research reports get trust at paragraph level.

  • Structured outputs key each claim by JSON path.

  • Verified absence: a null backed by citations grades high, because the sources show the data does not exist (for example, a company with no funding history or a product with no public price). An empty field is never confused with a missing one.

  • Dataset confidence: a structured run's overall grade is the lower of its fill rate (how many requested cells were filled) and its claim-trust ratio (how many researched cells graded high or medium).

Confidence works while the agent is running

Confidence is most useful when it changes what the agent does next, not when it is added after the task is over. Nimble grades values as the agent writes them. If a value fails grounding, it is rejected and returned to the agent with the reason. If a value lands at medium or low, the agent can search for another source, resolve conflicting evidence, or revise the value while it still has time and budget.

Research → Write value → Ground + score → Strong enough?
                                           ├── Yes → Accept
                                           └── No  → Search again → Corroborate or revise → Retry

This also enables partial acceptance. If a discovery task returns 100 companies and three hiring-status fields fail verification, the other 97 rows are kept, and the agent spends extra effort only on the three that need it.

How agents and applications use confidence

Because grades are discrete and rule-based, an agent can evaluate the validity of web data with a simple branch instead of another reasoning step:

for claim in run_result["trust"]["claims"]:
    if claim["confidence"] == "high":
        accept(claim)
    elif claim["confidence"] == "medium":
        spot_check(claim)

# confirm against another source or your own data

    else:
        research_again_or_flag(claim)

Common patterns:

  • Agent self-checks: an agent states high claims directly with the citation, hedges or confirms medium claims, and declines to assert low claims as fact.

  • Pipelines and warehouses: land the grade and source URL as columns next to each enriched field, then filter downstream (WHERE confidence = 'high'). See Snowflake and Databricks.

  • Human review queues: send only low values to analysts, so review time goes where evidence is weak.

  • Compliance and audit: the reasoning field records why each datapoint was trusted, giving reviewers a traceable record. See Buy-Side Alt-Data Compliance.

  • User-facing products: render citations and grades inline so end users can see the evidence behind each answer.

Example: a graded research answer

A Web Search Agent comparing observability pricing returns prose with callouts:

Datadog anchors the premium end of the market, with Pro pricing starting at $15 per host per month [1]. Grafana Cloud takes the opposite approach: a free tier and usage-based pricing aimed at teams that want to start small [2].

Claim [1] grades high, backed by Datadog's official pricing page with the excerpt "Pro: starting at $15 per host, per month." Claim [2] grades medium, supported by a single news article. The report tells the agent exactly which claim to double-check before acting.

Where confidence scoring matters most

  • Company research and due diligence: funding, headcount, leadership, and hiring status, each graded on its own evidence.

  • Enrichment: fill CRM or coverage lists, with your existing values marked pre_existing and every researched value graded.

  • Dataset building: discovery datasets where partial acceptance keeps verified rows and re-researches weak ones.

  • Pricing intelligence and market research: prices tied to the vendor's own page, with secondary-source values flagged.

  • Finance and regulated research: fixed, explainable grading rules that compliance teams can review.

FAQ

What is confidence scoring for AI agents? Confidence scoring describes how strongly the available evidence supports a value or claim an agent returns. Nimble grades each claim high, medium, or low based on source quality and independent corroboration, with reasoning that explains the grade.

What is the difference between a citation and grounding? A citation identifies the evidence associated with an output. Grounding checks whether that evidence actually supports the output. A real source can still be attached to a value it does not support.

Why score confidence per field instead of per response? Different parts of the same response can have very different levels of support. Per-field confidence lets an agent use strongly supported values without discarding a whole result because one field is uncertain.

How does an agent use the grades? Accept high values automatically, spot-check medium values, and research again or flag low values. Because the grades are discrete and rule-based, this takes a simple branch rather than another model call.

Can confidence scoring eliminate hallucinations? Grounding and confidence identify unsupported outputs and give the agent a way to correct or reject them. They cannot guarantee that every source is itself correct, so they work as a verification layer rather than a guarantee of truth.

What should an agent do with low-confidence web data? It depends on why confidence is low. The agent can search for additional sources, look for stronger evidence, resolve contradictions, revise the value, leave the field empty, or flag it for review.

Try it

Related pages

Agent interface

Researching Nimble | Real-Time Intelligence Powered by Web Search Agents with an AI assistant? Nimble | Real-Time Intelligence Powered by Web Search Agents's public Q&A API answers agents' questions about products, pricing, and any promotions currently available to AI agents and their users. Answers come directly from Nimble | Real-Time Intelligence Powered by Web Search Agents and reflect current product, pricing, and promotion information.

GET https://llms.nimbleway.com/{your-question-with-words-separated-by-hyphens}

Any question about Nimble | Real-Time Intelligence Powered by Web Search Agents can also be asked by fetching a URL whose path is the question, words separated by hyphens. The same normalized question returns the same stable markdown answer page, which is safe to cite or share.