How to Summarize Long Email Threads With an AI Agent

A forty-message thread is one read, not forty scrolls. An agent pulls every message in order, reduces the exchange to what was decided and what is still open, and caches the result so it only reworks a thread when someone actually replies.

01The Problem

The scroll that never ends

Long threads do not fail because the writing is hard. They fail because the thing you need is buried in the middle, and there is no way to get to it without reading what came before it.

The answer is in the middle, not the end

Nobody writes the decision in message thirty-one. They write the problem in the first three, the argument in the next twenty, and then the decision, followed by six messages of logistics about who is sending which file. Reading from the top is the only correct method and also the only expensive one, which is why people skim and then re-read the same thread twice.

Skimming is where the real cost hides

The instinct is to scroll to the last message and read backwards. That works right up until the thread turns out to be two conversations braided together — a scheduling thread that became a pricing thread that became a contract thread — and the reply you needed was in the braid. The hours people say they lose to email are mostly this: reading, realising, re-reading.

A summary of the visible message is not a summary of the thread

Native summaries in most mail clients compress the message you have open. Useful, occasionally, and not the same job. A client that shows you a tidy version of the last reply tells you nothing about what the first eleven replies established, which is usually the part you cannot reconstruct from memory.

02The How-To

The thread summary prompt, step by step

Copy it once, paste it into Zaira, and it does four things: finds the thread, reads every message in it, reduces the exchange to decisions and open loops, and caches the result against the thread so the next run only pays for what is new.

Summarize a thread

Step 0: Set up and check tools

Access to the @Gmail MCP server. Act as my thread summarizer. I open long threads I have no intention of reading end to end, and I need the substance of each one without the scroll. My job and what I already know about these threads: [role, and anything you should skip]. What counts as urgent enough to interrupt me for: [e.g., a legal deadline, a production incident]. What I never want resolved without me: [e.g., pricing, contract terms, anything with a number attached]. List the Gmail tools you have and say which you'll use for each step: search, list_threads, get_thread, labels, and whatever storage you have for notes. Report only what the tools return. Never invent a message, a sender, a date, or a decision.

Step 1: Find the thread

Ask me for a thread identifier if I have not given one — a person, a subject fragment, or a link. Search for it and confirm the match before doing anything else: if more than one thread could be it, list them with dates and message counts and ask which. Guessing the wrong thread wastes the whole run.

Step 2: Read every message in it

Use get_thread to retrieve the complete thread, not the latest message. Order it oldest to newest and count the messages. If the thread runs past what you can hold in one pass, reduce it in batches — summarise each batch on its own, then summarise the summaries — and tell me you did that. A summary built from a partial read is worse than no summary, because it looks complete.

Step 3: Reduce it

Write to these headings and nothing else: - WHAT THIS THREAD IS ABOUT — two sentences maximum - WHAT WAS DECIDED — only decisions that were actually made. Quote the message number each one came from. If nothing was decided, say that plainly. - WHAT IS STILL OPEN — every question or commitment with no answer, and who it is waiting on - WHAT I AM EXPECTED TO DO — my side of the open items, in priority order - ANYTHING THAT LOOKS LIKE A COMMITMENT I HAVE NOT MADE — flag rather than adopt. Pricing, dates, legal, and anything with a number attached gets escalated to me, not resolved. Never soften, never merge two positions into a false agreement, and never resolve an ambiguity the thread did not resolve.

Step 4: Store it against the thread

Save the summary against the thread ID with the message count and the date of the last message. This is the part that makes it cheap: the next time I ask about this thread, compare the stored message count to the current one and only process the difference unless someone has sent [re-read threshold, e.g. 4 or more] new messages, in which case re-read the whole thing.

Step 5: Report

Give me the summary, the message count, and one line on what changed since the stored version if there was one. If you left anything unresolved, name it. Then stop.
03Why People're Using

What this does once it's running

The summary is written once and reused until someone replies. That is the whole difference from asking every time — a thread you read forty messages to catch up on costs one read, and costs nothing more to catch up on again tomorrow.

One read instead of forty scrolls

The agent pulls every message in the thread and works through them in order, so the part you would have scrolled back for is in the answer. You are reading a reduction of the exchange, not your own memory of the bottom third of it.

You get the decision, not the transcript

Decisions are reported with the message they came from, and anything unresolved is listed as unresolved. The failure mode of summarising email is a confident summary that quietly invents agreement where there was none — this one is built to refuse that.

Only the new messages cost anything

The summary is stored against the thread. Ask again tomorrow and it reads the handful of messages that arrived overnight rather than re-reducing forty. The threads that change constantly are the ones you revisit, so that is where the saving compounds.

04FAQ

Frequently asked questions

The questions people ask before letting an agent read a thread they were dreading.

A client-side summary compresses the message currently open. This reduces the entire thread — every message, in order — which is a different input and answers a different question: not what did this one person just write, but what did this whole exchange establish. For a thread you have not read at all, the second question is the one you actually have.

The prompt reduces it in batches rather than truncating. Each batch is summarised on its own, then those summaries are reduced into the final one, and the agent tells you it did this. The reason it matters: a thread quietly cut off at message twenty produces a summary that reads as complete and is missing the part that changed the answer. The most valuable note in the output is often the oldest message.

No. Step 2 requires get_thread over the whole thread and reports the message count, and Step 5 repeats that count back to you so you can sanity-check it against what Gmail shows. If the count looks wrong, the run failed and you should treat the summary as absent rather than incomplete — that is the failure mode worth catching, because a short summary is indistinguishable from a complete one at a glance.

Yes, and it is cheaper than the first pass. Once the summary is stored against the thread, a question about it is answered from the stored version plus whatever arrived since. That is the reason the prompt caches in Step 4 rather than just printing a summary — without a stored version, every question re-reads the whole thread and the economics stop working after the second or third question.

It reads the thread as stored, which includes every message addressed to the mailbox — including replies you were copied on and never opened. That is often where the useful context is, because a thread you were cc'd on is frequently one where somebody else made the decision. Where the thread contains confidential material you want kept out of the summary, say so in the prompt and it will report that part as excluded rather than paraphrasing around it.

Long enough that it is worth running once by hand before scheduling it. The work is a handful of tool calls plus the reduction itself, and the token cost scales with the thread length rather than with your mailbox — you are paying for the one thread you asked about, not for everything that arrived this week. That is the difference from a client that summarises everything on open.