How to Run Deep Web Research with an AI Agent That Cites Sources

A prompt that turns a vague question into a real research brief — several search angles, a reading list, and a structured answer where every claim carries the link it came from and disagreement between sources is reported rather than smoothed over.

01The Problem

Why one search is never enough

The first page of results is an advertisement for what other people searched for, and the interesting answer is usually on page four.

The first page is optimised for clicks

Ranking reflects what has been linked to and paid for, not what answers your question. The page you want is frequently not in the top ten, and the confident piece at number one is often a vendor explaining why their approach is the right one. A single search takes the shape of what other people were already looking for, which is the one thing you were trying to get away from.

Reading is the expensive part, not finding

Twenty tabs open and none of them read properly is a familiar afternoon. The bottleneck is comprehension — working out which of these sources are independent, which are citing each other, and which are old — and that is the work that does not survive being skimmed. Which is why the finding is often better than the searching, and why more links rarely help.

Sources that agree are not sources that confirm

When five pages say the same thing, it is usually because they all read the same study, and treating that as five confirmations is how confident wrong answers get written up. Nothing in the reading experience flags it. Establishing whether two sources are actually independent takes a specific check that almost nobody performs, and it is the difference between a brief you can rely on and one that merely looks well-sourced.

02The How-To

The web research AI agent prompt, step by step

Copy it once, paste it into Zaira, and it runs the research properly: builds several search angles, fetches and reads the actual pages, cross-checks the claims that matter, and hands back a brief with sources attached to every line.

Deep web research
Access to the @Exa MCP server. Act as my research agent. My goal is a sourced brief on a question I could not answer properly myself in an afternoon — built from pages that were actually read, with every claim carrying the link it came from.

The question

- What I actually need to decide: [the real question, not the topic] - Why it matters and what I will do with the answer: [e.g., deciding our pricing model for Q2] - Time horizon the answer needs to be current to: [e.g., last 6 months] - Depth and budget: [e.g., thorough, can spend a minute of your time; or quick, one pass] - Anything I already know or suspect, so you do not repeat it: [notes]

Step 0: Check tools

List the Exa tools you have (web search, advanced search, web fetch, agent run) and say which you'll use for each step. Only report what you actually retrieved — never cite a URL you did not fetch, never cite from memory, and never write a source line for a page you did not open.

Step 1: Plan the angles, not just the query

Before searching, break my question into 4-7 angles a competent researcher would take: the primary sources rather than the summaries, the counter-arguments, the recent developments, the practitioner experience versus the vendor claims, and the numbers. Show me the angle list with the search query for each. Then search all of them.

Step 2: Read, do not skim

Fetch and actually read the most promising pages from each angle — not the snippets. For each one, record: what it claims, what evidence it gives, when it was published, who benefits from the claim, and whether it cites its own sources. Skip anything that is a vendor explaining its own product, and say so when you skip it.

Step 3: Check independence

For every claim that matters to my decision, trace where it came from. If several sources agree, determine whether they are independent or all repeating the same underlying study, report, or press release. Flag every claim resting on a single originating source, and any claim where sources actively disagree — do not average disagreement into a confident middle position.

Step 4: Write the brief

Structure it as: the direct answer in three sentences; the evidence for it, with a link and date on every claim; what the sources disagree about; what nobody seems to have answered; and what would change the answer. Separate what the sources say from your own inference. Every factual sentence carries its source inline, and anything you could not verify is marked as unverified rather than omitted.

Step 5: Report

Give me the reading list with one line on why each source mattered, the claims that rest on a single source, and the two things most worth a follow-up question. Keep it under 250 words — I will read the brief, not your notes.
03Why People're Using

What this does once it's running

Three stages, one goal. Questions that used to mean an afternoon of tab-management produce a sourced brief you can check in minutes — and re-run cheaply when the answer is six months stale.

Several angles instead of one query

The plan in Step 1 is what makes this research rather than a search: primary sources over summaries, counter-arguments, recent developments, and practitioner experience instead of vendor claims. The answer is usually in an angle nobody searched from.

Every claim carries its link

Inline sourcing on every factual sentence, plus an instruction never to cite a page that was not actually opened. A brief you can verify claim by claim in a minute is worth more than a longer one you have to trust.

Disagreement gets reported, not averaged

Where sources conflict the brief says so, and where several sources trace back to one study it says that too. That independence check is what stops a widely-repeated claim being presented as a widely-confirmed one.

03bField Notes

What actually happened when we ran this

One real run of Steps 1 to 5 above, written up as it happened. The retrieval backend was our own search tooling rather than Exa, so treat the angles as specific to this run rather than to the method.

Step 3 caught two papers that looked like corroboration and were the same measurement

Searching for citation-accuracy research returned Onweller et al., 'Cited but Not Verified' (arXiv 2605.06635, 7 May 2026), reporting that Fact Check accuracy drops ~42% as tool calls scale from 2 to 150. A second hit, Leung et al. (arXiv 2607.08700, 9 Jul 2026), independently states the same ~42% figure — which reads like two labs confirming each other. Fetching it shows the opposite: it says outright that it builds on the pipeline of Onweller et al., and three of its four authors (Lumer, Huber, Feld) are on both papers. One measurement, cited twice. Without the independence check in Step 3 this would have been written up as convergent evidence, and it is the single most valuable thing the step caught.

The counterintuitive result was that deeper research made citations worse

The instinct behind a research agent is that more retrieval means better grounding. The measurement points the other way: factual support degrades as the agent reads more, across the 2-to-150 tool-call range. That single finding reframed the whole run — it moved the emphasis off 'find the right sources' and onto 'check whether the source says what the claim says', which is what Steps 2 and 3 actually enforce. Worth noting the same paper reports link validity above 94% and relevance above 80% while factual accuracy sits at only 39-77%: the metrics a reader can eyeball are the healthy ones.

Two of the five most promising results were unusable, and neither was obvious from the title

The most eye-catching results were a Perplexity-published benchmark and a vendor-adjacent leaderboard. DRACO (arXiv 2602.11685) is hosted at hf.co/datasets/perplexity-ai/draco and reports Perplexity Deep Research ahead of every model tested, while being built from that company's own production tasks. Step 2's instruction to note who benefits from a claim is what flagged it. Both were demoted to context and excluded from load-bearing claims. Neither advertises its own conflict in the title, so the exclusion had to come from the affiliation rather than the abstract.

The most useful numbers came from the two results ranked lowest for relevance

By the search engine's own ordering, the two papers that carried the actual figures — AutoResearchBench at 9.39% Deep Research accuracy and the AARR benchmark at 68.3% for the best configuration — sat below the vendor and aggregator pages. Had Step 1 stopped at 'what is already published on this topic' rather than splitting it into primary measurement versus summary, the run would have reported confident numbers from press coverage and missed that frontier research agents score below 10% on hard literature-discovery tasks. That gap is the argument for the multi-angle plan in one concrete instance.

04FAQ

Frequently asked questions

The practical questions people ask before trusting research an agent did for you.

The instruction is absolute in Step 0 — never cite a URL that was not fetched and never write a source line for a page that was not opened — and it is enforced by the structure rather than by trust, because Step 2 requires a recorded claim, date, and evidence per page before anything gets written. The quick verification is the brief itself: click three sources. If they say what the brief claims, the citation pattern holds; if one does not, you have found the limit of the setup immediately.

A single answer from a model's memory has no reading list, no dates, and no way to check whether five agreeing sources are one study quoted five times. This is built around the opposite: plan several search angles, fetch and read the pages, trace claims back to their origin, and report disagreement. It costs more time per question, which is the trade — you are buying a brief you can verify rather than a paragraph you have to trust.

Questions with a discoverable answer that needs breadth — competitive landscape, market sizing, how a regulation affects a specific case, what changed recently in a field you follow, the state of a tooling category. It is much weaker at anything requiring primary data nobody has published, at precise figures that need a paid database, and at anything where you need to talk to someone. Worth being honest about that, because it is where every research agent's ceiling is.

It depends on how many angles Step 1 produces and how many pages Step 2 reads — that is the honest answer, and the depth setting in your context section is the dial. Zaira charges per plan rather than per task, so a broader sweep is not a separate bill. Because it reads pages rather than snippets, a thorough run is genuinely more expensive in time than a quick one, which is why the prompt asks you to say up front how deep this needs to be.

It works only with what Exa can retrieve and reads public pages, so anything behind a paywall or a login is simply not available to it — and the prompt is told to report a source it could not read rather than paraphrase from a snippet. If your research depends on paid sources, fetch them yourself and put the extracts in your context section; the agent will work from what you give it and cite your material rather than a page it could not see.

Yes, and it is one of the better uses of it — set the time horizon in your context section and schedule the same prompt to re-run monthly or quarterly. Because Step 3 traces claims to their originating source, a re-run surfaces a changed figure and points at where the new number came from, rather than quietly restating the old one as current. Topics with dated numbers are where this earns its keep fastest.