The Unpaid Invoice: Meta Trains. You Pay Per Request.
In brief
Meta's crawler hit steampunkfarms.org 12.5 million times in eighteen days. We paid $88.75 for the privilege of being copied. Zero donors came back the other way. This is not a complaint about how the internet works. It is an invoice, itemized, with a name on it — and an argument that the cost belongs on Meta's balance sheet, not ours.
The Unpaid Invoice: Meta Trains. You Pay Per Request.
Seven hundred thousand requests a day. On the worst day, a million and a half.
Not readers and not donors. Not one person who might send twenty dollars after watching a lame goat get her feet back under her. A verified Meta crawler, hitting steampunkfarms.org hard enough to take that site's hosting line from fourteen cents a day to $12.88.
I want to be exact about this, because exactness is the only thing that survives contact with a company that size. What follows is off the invoice, itemized, with a before and an after.
Between August 14 and August 31, 2026, the logs on steampunkfarms.org recorded 12,576,223 edge requests. In a sampled ninety-minute window, 93.3 percent of all traffic to the site came from AS32934, which is Facebook, Inc., from 223 distinct addresses inside a single 57.141.0.0/16 block Meta owns outright. Every one of them identified as meta-externalagent/1.1 and passed the host's verified-bot check. Nobody was spoofing anybody. This was Meta, crawling as Meta, at scale.
Human traffic over the same window: flat. Donations: flat.
The hosting line for that project went from $0.14 a day to $7.97 a day across the peak week, and $12.88 on August 22 alone. Eighteen days of crawl cost $88.75. Run the peak-week rate out and it is about $235 a month, or $2,859 a year, on a sanctuary's books. Those two are extrapolations from a measured daily rate, and I will say so every time I use them; we have not been crawled for a year, only for eighteen days — and the eighteen days were enough.
They copy the site to feed a commercial model. We pay the meter that counts the copying. That is not crawling. That is an invoice with a friendly user-agent string and no return address.
The receipt
I got one thing wrong before I pulled the billing export, and the correction makes the case worse for them, not better for me.
I assumed we were paying for bandwidth. We were not. Bandwidth cost us nothing. Fast Data Transfer for the whole crawl came to 144 gigabytes against a 1,000 gigabyte allowance: $0.00 billed. The crawler even negotiated compression like a good citizen, pulling about 17 KB a request instead of the 1.6 MB the uncompressed page would have cost.
The bill was the request itself, metered four separate ways.
Line item | Volume, Aug 14–31 | Billed |
|---|---|---|
Observability Events | 28,356,034 | $32.68 |
Edge Requests | 12,576,223 | $15.91 |
Function Invocations | 13,094,835 | $7.56 |
ISR Writes | 1,335,166 | $5.09 |
Fast Data Transfer | 144 GB | $0.00 |
That is the modern hosting stack, and anyone who wants to argue about "the cost of being online" should read the table first. Nobody bills you for the pipe anymore. They bill you for the event. One crawler hit is an edge request, and an observability event, and, if it misses cache, a function invocation and a cache write. Four meters, one visitor who was never a visitor.
The perverse consequence is that a polite crawler costs you what a rude one does, to within a rounding error. Meta compressed its downloads. It saved us nothing, because nothing we pay for is measured in bytes.
The single largest line, $32.68 of it, the biggest cost on the whole site, is observability. The instrument that tells you a bot is eating your site is itself billed per event, so the bot inflates the very meter you would use to find it. We turned observability off on August 29 to stop the bleeding. In the billing data that decision is a column of zeros, and what the zeros mean is that for three days we were being crawled blind. That is the shape of the trap: the crawl makes the diagnosis expensive, so you switch off the diagnosis, so the crawl continues unobserved... and the invoice keeps arriving either way.
Before August 14 this site cost fourteen cents a day to run. It is a sanctuary website. It is supposed to cost fourteen cents a day to run.
The rate limit that caught nothing
We tried the obvious thing first. A per-address rate limit, set well above anything a human could trigger.
Zero denials. Not one, across hundreds of thousands of requests.
That is the 223 addresses doing their work. No single address ever went fast enough to trip a per-address limit, because the load was spread across a /16 that Meta owns. Every small-site defense on every forum, block the IP, throttle the IP, is aimed at a shape of attacker this is not. You cannot rate-limit a company that has more addresses than you have visitors. Our firewall did not fail. It measured the asymmetry and reported it correctly.
You are not crazy; the fleet got fatter
The temptation when this happens is to assume you misconfigured something. You did not. The fleet grew this year, and the one industry number I trust enough to print says so.
DataDome's second-quarter report for 2026 puts Meta-ExternalAgent at 5.3 billion requests across their network, up 74 percent from the quarter before, and meta-webindexer at 3.75 billion, up 163 percent. Nine point one billion requests in ninety days from one company, and the report's own phrasing for what came back the other way was "almost no real visitors." That is the whale. The rest of the pod is real: OpenAI, Anthropic, ByteDance, Amazon, Perplexity, and a long tail of "research" agents with no contact address. On our logs the whale is Meta at 93.3 percent, with the rest of the fleet sharing the remaining seven, so I am naming the whale.
Read the Docs, which hosts documentation for open-source projects on donated money, wrote in 2024 that a single crawler had pulled 73 terabytes from them in a month. Blocking AI crawlers cut their daily bandwidth by about three-quarters and saved them on the order of $1,500 a month. They were paying for bytes. Two years later we are paying for events, and the meter got smarter faster than the defenses did.
Large media companies license. Reuters, Le Monde, CNN, Fox, USA Today have deals with the labs, terms undisclosed. Small sites donate. That is the class structure of the training economy, and nobody voted on it.
The fee with no name, again
In July I wrote about interchange: the 2.36 percent that Visa and Mastercard skim from every card swipe, hidden in the shelf price so the cash customer subsidizes the rewards customer and nobody sees the transfer. The argument was simple. Operating costs belong in the extractor's pricing, not on someone else's receipt.
Line the two up.
The old skim | The new skim |
|---|---|
Visa and Mastercard interchange | AI training crawl and answer index |
Merchant eats 2.36 percent | Site owner eats a per-request meter |
Cash customer subsidizes card rewards | Nonprofit subsidizes a model that will never send a donor |
Fee hidden in the shelf price | Crawl hidden in "the open web" |
The same shape and same move. A cost that belongs on the balance sheet of a company worth more than a trillion dollars lands on the statement of a rescue barn with 88 goats and a mortgage.
A surcharge at checkout is rude. Seven hundred thousand daily hits with no customer behind them is a surcharge you never agreed to and cannot decline at the register, because there is no register. There is only your server, answering politely, and a bill at the end of the month.
What they buy with your invoice
Skip the legal theory. This is a character question.
They buy tokens. Every page on steampunkfarms.org that a volunteer wrote at eleven at night about a pig named Penelope becomes a few hundred tokens in a corpus, and the corpus becomes a model, and the model gets given away so Meta can own the surface people ask questions on. Sam Altman said out loud at BlackRock that the plan is to own the pipes and charge for intelligence by the unit. I wrote that one up too. Meta's version skips the meter on the front end and runs it on ours. Either way the corpus is the raw material, and the raw material was collected at the supplier's expense.
They also buy an index, which is the sneaky thief. Search, whatever its sins, at least pretended to send you a visitor. The deal was always crooked, but it was a deal: you let Googlebot in, Googlebot sent people. The answer engines do not pretend. They read the whole site so they can tell someone what the site says without that someone ever arriving. The donor asks Meta AI whether the sanctuary takes surrendered roosters. Meta AI says yes. The donor never sees the donate button. We paid for the request that taught Meta the answer.
The number that ended the argument in my own head: over that window Meta made twelve and a half million requests to our site and sent us no measurable traffic at all. A denominator with nothing over it.
"Publicly available" is doing a lot of work in the industry's talking points, so let me take the load off it. Public means visible. A barn on a county road is visible. That does not license a trucking company to run a convoy through the driveway on my diesel because the gate was open. Visible does not mean free to industrialize at the publisher's expense, and it never did. The word only started meaning that when the people who wanted it to mean that had lawyers.
The architecture of the threat
Who it harms tells you who it serves.
The winners at scale are the labs, and behind them the publishers big enough to have a licensing desk and a general counsel who can get a Meta lawyer on the phone. A ninety-dollar overage on a Fortune 50 property is dust on a line item. Somebody in procurement would have to squint.
The losers at scale are the WordPress church, the sanctuary site, the indie documentation project, the one-woman shop on a $12 host plan, the hundred thousand sites that exist because a person cared about something and learned enough HTML to say so. Every one of them is being crawled at the same intensity as the New York Times, because the crawler does not grade on size. The bill lands on the entity least able to absorb it, and the entity least able to absorb it has the fewest ways to say no.
On a sanctuary ledger $235 a month has a name, and the name is not "infrastructure." It is 8 bags of pig feed. It is the diesel for 5 emergency hay runs. It is the float you keep for a vet call so you do not have to decide between the vet and the grain. It is the matching gift you already spent a Saturday on the phone begging for, gone, before the animals saw a dollar of it. Meta's corpus or the pigs' grain. Same dollars, and only one of the two parties gets to choose.
If the site slows down, donors bounce. If you take the site down to stop the bleeding, the animals do not get a press release. That is the whole asymmetry. They can afford to be indifferent. We cannot afford to be crawled.
What a steward does
This is not a tutorial. This is the list of doors, with the cost of each door named, because a steward does not get to pretend the exits are free.
The first door is robots.txt, with the AI crawlers named and disallowed. It costs nothing and it is the honorable request. The catch is that it is a request. File it anyway, because the record will count if this ever reaches a courtroom or a regulator. Do not mistake it for a lock.
Know which agent you are naming, too. "Block Meta" is five switches, and one of them is wired to your own house. meta-externalagent trains the model. meta-webindexer feeds Meta AI's citations. meta-externalads improves their ad products. facebookexternalhit is the crawler that builds the link-preview card when you post your own fundraiser to Facebook or Instagram: the title, the description, the photo of the goat. Block that one by accident and you have not stopped Meta from taking. You have only stopped your own posts from working. It is the token most people reach for first, and it is the wrong one.
The second door is the firewall. It works right up until they rotate addresses, which they do — 223 of them in two and a half hours, out of a block they own. Reach for an ASN-wide block to catch them all and you take your own link previews down with it, because all five crawlers share AS32934. The precise rule is a user-agent rule. The blunt one costs you the thing you were protecting.
The third door is an edge cache, so the crawler hits a static copy and your origin never wakes up. Do this regardless. It is the difference between a donor waiting four seconds and one second. Understand what it does not do. On a per-request meter a cached hit is still a billable hit. Caching made our pages 80 percent smaller and it did not zero the bill, because the bill was never about size.
The fourth door is pay-per-crawl, the HTTP 402 experiments where the site names a price and the crawler pays it or leaves. I like this door because it does the one thing I have been asking for since July. It names the invoice. Its catch is leverage: a lab can walk away from any single site, and a rescue barn negotiating alone with Meta is a mouse negotiating with a combine. This door only works if enough small sites stand in it at once.
The fifth door is architecture. Starve the origin: static files, no database, nothing expensive to hit. We did a version of this. We paginated the heaviest page, moved its filters into pre-rendered paths, cut it from 1.67 MB to 345 KB. It was worth doing and it helped real people on real phones. It did not stop the meter. I wrote in The Cheapest Governance that architecture is also a bill — this is the case where the cheapest architecture still leaves you paying for someone else's appetite.
The sixth door costs nothing and almost nobody walks through it: refuse the framing. Publish the bill. Say the number out loud. Make "training data" a political noun again instead of a technical one. Every site owner who says "well, that's the cost of being online" has already paid and agreed to keep paying. Every one who prints the invoice moves the cost one inch back toward the company that incurred it.
You do not have to win against Meta. I wrote in The Mathematics of Return that the quarter-measure is the whole game: take back a quarter of what the machine does to you and you are ahead of everyone who took back nothing. You have to stop paying in silence. That is all. Half the exits above are cheap. The last one is free.
The close
They will call this the cost of being online.
It is the cost of them being online on your meter, and the two are not the same thing no matter how many times a press release says "open web" in a tone of voice that implies you should be grateful.
A household that will not name an invoice will keep paying it. We are naming it.
steampunkfarms.org. August 14 to August 31, 2026. 12,576,223 requests. 223 addresses. $88.75 in eighteen days, $2,859 a year at the peak rate. Zero donors. One whale.
Send the bill to Menlo Park. I already know they will not open it. The point was never that they would.
Coda: The Second Bill
I have been here before with this company, and the two halves only make sense together.
Before the crawler, we raised real money on Facebook. Thousands a month, from people who followed us because they wanted to see the animals. That was the arrangement, and it was a good one. We made something worth following, people followed it, and when we asked for help the people who had chosen to hear us heard us.
Then the reach went away. Never all at once, never with an announcement. A post that had reached thousands of our own followers reached hundreds. Then dozens. The fundraisers we had run every year, on the same day, to the same list, started coming back with a fraction of what they used to. Inside of twelve months the thing that had paid for feed was paying for almost nothing.
There was always an option to fix it, sitting right there in the interface. Boost this post. Reach the people who already asked to hear from you, for a fee.
I have heard the counter-arguments. More content, finite feed, ranking is hard, it is not personal. I know it is not personal. That is rather the point: indifference and extraction are indistinguishable from underneath. What I know is what every small nonprofit I talk to knows, because we all watched the same graph do the same thing in the same period. We built the audience, on their platform, at their invitation, and then we were charged rent on our own followers.
That is the first bill. The second one lands differently.
Having throttled the channel we used to raise money, they are now spending our money to take the thing we made instead.
Look at the shape of it. We wrote every page on that site. The intake notes, the surrender-prevention guides, the transport protocols, the vetting diagnostics, the years of learning how to get a downed goat back on her feet — all of it written by hand, at night, after chores, by people who were not paid to write it. Meta's crawler took twelve and a half million passes at it in eighteen days, and we paid, out of the donations that were supposed to buy grain, for the privilege of being copied.
They reduced our ability to raise money. Then they billed us for the raw material.
I do not think anyone at Menlo Park decided to do this to us. Nobody there knows we exist. That is what makes it a structure rather than a grievance, and structures are the only things worth writing about, because a grievance ends when you feel better and a structure keeps running.
The structure is this: the same company that took the distribution is now taking the content, and we are the ones being metered for both transactions. First they made our audience harder to reach unless we paid. Now they make our writing more expensive to host while they train on it. One of those is a business model. Both of them together is a toll road built on a road we paved.
I am careful about the word "stealing," because it will be the first thing someone reaches for to dismiss this. I am not making a copyright argument. I am making an accounting argument, and accounting is harder to argue with. Whatever the law eventually decides about training data, the electricity bill is not in dispute. Somebody paid to serve those twelve million requests. It was not Meta. It was a 501(c)(3) with pigs and goats and cows and cats.
We shut the crawlers out. Not the one that makes our own posts work; we still need that, because we still, begrudgingly, need the monster. Just the ones that take without sending anyone. It took an afternoon and it will not survive contact with a company that decides to ignore it.
The ledger is straight now, and that was always the actual goal. When Meta comes back, and they will, under a new user-agent string, the way they always do, the answer will already be written down, with a number attached.
We paid twice. We are done paying in silence.
— F. Tronboll III
Common questions
How much did Meta's crawler cost a small sanctuary website?
$88.75 in eighteen days, with a peak of $12.88 on a single day. At the peak-week rate, that extrapolates to roughly $2,859 a year. The site had been costing fourteen cents a day before the crawl began.
What exactly was billed — bandwidth or something else?
Not bandwidth. Fast Data Transfer came to $0.00. The bill was the request itself, metered four separate ways: edge requests, observability events, function invocations, and ISR writes. Nobody bills for the pipe anymore; they bill for the event.
Why didn't rate limiting or IP blocking stop the crawl?
Meta spread its load across 223 distinct addresses inside a single /16 block it owns outright. No single address ever went fast enough to trip a per-address limit. You cannot rate-limit a company that has more addresses than you have visitors.
What is the difference between Meta's crawler types, and which ones should a small site block?
meta-externalagent trains the model. meta-webindexer feeds Meta AI's citations. meta-externalads improves ad products. facebookexternalhit builds the link-preview card when you post your own fundraiser — block that one and your own posts stop working. The precise rule is a user-agent rule; a blunt ASN block costs you the thing you were protecting.
Did edge caching reduce the bill from AI crawlers?
Caching made the pages smaller but did not zero the bill, because the bill was never about size. On a per-request meter, a cached hit is still a billable hit.
Is this a copyright argument or something else?
It is an accounting argument. Whatever the law eventually decides about training data, the electricity bill is not in dispute. Somebody paid to serve those twelve million requests. It was not Meta.
Takeaways
- Modern hosting meters events, not bytes — so a polite AI crawler costs the site owner the same as a rude one, because the bill is per request, not per gigabyte.
- The crawl inflates the very observability meter you would use to detect it, creating a trap: diagnose the problem and pay more, or switch off the diagnosis and go blind.
- Large publishers license their content; small sites donate it involuntarily — that is the class structure of the training economy, and nobody voted on it.
- Robots.txt is an honorable request, not a lock — file it anyway because the record will count if this ever reaches a courtroom or a regulator.
- Refusing the framing costs nothing: publish the bill, say the number out loud, and make 'training data' a political noun again instead of a technical one.
F. Tronboll III
Share