How to Save Email Attachments to Google Drive With AI

Filing by filename puts every invoice in a folder called documents and every contract in one called final2. An agent opens each attachment, works out what it actually is, and puts it somewhere findable — by vendor, type and date.

01The Problem

Two hundred files called document.pdf

Nobody fails at saving attachments. They fail at naming them, and then the download folder becomes the archive, and the one thing the filing was for — finding last quarter's AWS invoice in February — becomes a search you are not sure you can win.

Filenames lie, and you cannot file on a lie

The attachment is called invoice.pdf or final-v3.docx or a 2MB string of characters from a scanner. Nothing in the name says which vendor, which month, or which client, so any rule you write against filenames either does nothing or files everything somewhere useless. This is why the useful automation reads the file rather than the name — and why tools that promise to sort by filename are quietly not solving the problem.

The archive and the inbox end up as the same place

An attachment left in Gmail is findable exactly once, while you remember which message it came from. The moment it matters is eighteen months later when you need the invoice and all you have is a search box that has never successfully found a PDF. Saving to Drive does not fix that on its own — saving it under a name that means something does.

The duplicates are already there

Most people assume they would recognise duplicates if they had a good system. They would not: the same invoice arrives as a PDF in March and again in April as a forwarded copy with a different subject line, and both are filed under different names in different places. A filing pass that reads contents catches those. One that reads filenames never will, because to it they are two different files.

02The How-To

The attachment filing prompt, step by step

Copy it once, paste it into Zaira, and it works the mailbox properly: find every message with an attachment, open the file rather than trusting its name, pull out what identifies it, file it under a naming convention you wrote, and report every decision it made.

File attachments

Step 0: Set up and check tools

Access to the @Gmail MCP server and to Google Drive. Act as my attachment filing agent. The job is to read each attachment, decide what it is, and put it somewhere I will find it in a year's time. My naming convention and folder structure: [e.g. Finance/Invoices/VENDOR/YYYY-MM-vendor.pdf; Clients/NAME/DATE-title.pdf]. Documents I must never touch: [e.g. anything from legal, anything with signed in the subject]. What counts as a duplicate for my purposes: [e.g. same vendor, same amount, same month]. List the tools you have for searching mail, reading attachments, creating folders, uploading files, and moving or renaming. Report only what they return. Never invent a filename, a folder, a vendor or an amount.

Step 1: Find the attachments

Search the mailbox for messages with attachments over the window I give you — [e.g. the last 12 months] — using Gmail search operators, and run several passes rather than one: newer_than, older_than, and has:attachment combined with a from: filter for the senders that matter most. Do not stop at the first page of results. Report how many messages you found before you file anything, and if it is a much larger number than you expected, say so and let me confirm.

Step 2: Read the file, not the name

For each attachment, open it and extract what actually identifies it: document type, vendor or organisation, date, and for an invoice the amount. Do this from the content of the file. If a PDF is a scanned image rather than selectable text, say so for that specific file and put it in an Unsorted list rather than guessing from the surrounding email. If you cannot tell what a file is from its contents, it goes to Unsorted. A wrong folder is worse than no folder, because it will never be looked at again.

Step 3: Check for duplicates before filing

Compare each attachment against what you have already filed and against what you have seen in this run, using the duplicate rule I gave you. If it is a duplicate, do not upload a second copy — report the original's path instead and move on. Count these at the end; on a long history the count is usually the most interesting number in the report.

Step 4: File and rename

Create folders as needed rather than assuming they exist. Upload each file under the name my convention produces, derived from the content you extracted in Step 2. Never overwrite an existing file: if a name is taken by something different, append a discriminator rather than replacing it. Leave the original in Gmail untouched — I want the mail to stay where it is, not to be emptied.

Step 5: Report

Give me: files processed, folders created, files per top-level folder, duplicates skipped, files sent to Unsorted with one line on each saying why, and anything that looked like an invoice but was not one. Then list anything you deliberately did not touch. Stop there — do not delete or move anything in Gmail.
03Why People're Using

What this does once it's running

Tax season becomes opening a folder instead of searching a mailbox. The filing rules are written once, in plain English, and the interesting part is the first pass over years of history — which is where you find out you have four copies of the same invoice under three names.

The folder structure becomes the archive

Once invoices land under their vendor and month, tax season is opening a folder. The search you were dreading becomes a folder listing, which is the whole reason anyone files anything at all.

Names that mean something

A file called 2024-03-aws-invoice.pdf is findable by search, by folder, and by eye. A file called invoice.pdf is findable only if you already know which of the two hundred invoices you want.

The duplicates get caught

Reading contents means the same invoice arriving twice is recognised as one invoice, however it was named. Filename-based filing cannot do this, which is why duplicate counts tend to be the first surprise people get.

04FAQ

Frequently asked questions

The practical questions people ask before handing an agent their filing system.

A rule does that, and it is the right first step. What it cannot do is name anything: it moves the file and leaves it called invoice.pdf, in one flat folder, forever. The value here is not the moving, it is the reading — pulling the vendor and date out of the file so the folder it lands in actually tells you something later. If you only want the moving, a filter does it in thirty seconds and this is the wrong tool.

It will tell you it cannot, and the file goes to an Unsorted folder with a note. That is deliberate rather than a limitation to work around: a scan of a receipt photographed on a phone can often be read, but a scan of a fax can usually only be guessed at, and a guessed vendor name in a financial folder is a problem you will not discover until someone relies on it. The report tells you how many landed there so you can decide whether to widen it.

Nothing. The prompt copies into Drive and explicitly leaves the original in the mailbox — no delete, no move, no label change. That is worth being clear about because filing tools that empty the inbox as a side effect are a common and unpleasant surprise, and a copy is what you want: the mail is your record of what was sent and when, and the Drive copy is the archive.

Long enough that you should not watch it, and long enough that the first run is worth scoping to a single document type. Run it once for invoices only, check the results by hand, adjust the naming convention, and then let it widen. Backfilling years of attachments in one pass is how you end up with four hundred files in Unsorted and no idea whether the agent was wrong or the convention was.

Give the existing structure to the prompt and it will create only what is missing and file into what exists. Giving it a folder you made is much better than letting it invent one, because the invention will be reasonable and therefore never revisited. If your structure is three levels deep with dates in folder names rather than filenames, say that too — that is a different convention and the prompt will follow it.

Yes, and that is the version worth having. Run it on a schedule against recent mail only, and the cost each run is proportional to what arrived rather than to how much history you have. The backlog pass is a one-off; the ongoing pass is what stops the download folder from starting again. Run the same prompt with a short window and you have both.