What share of web traffic is generated by bots?

We picture the web as a space full of humans who click, read and buy. The reality is stranger: in 2024, for the first time in ten years of measurements, automated traffic overtook human traffic. In other words, one in two requests reaching a site is no longer sent by a person, but by a program.
This figure deserves a pause, because it changes how we understand what happens behind a simple page view. Here is what the recent data says, where it comes from, and why it can't all be read the same way.
Not all bots are equal
A bot is simply software that interacts with a site without direct human intervention. The category covers the best and the worst.
On one side, the good bots: Google's crawler that indexes pages for search, uptime monitors, legitimate price aggregators, tools that check links. They identify themselves, respect the rules published in the robots.txt file, and don't try to hide.
On the other, the bad bots: those that try to guess passwords, buy out an entire stock of tickets in a second, copy a whole catalog, or skew advertising statistics. These are the ones that raise concern, and their share keeps growing.
The trend is clear and steady: five consecutive years of increase. Imperva attributes part of this acceleration to generative artificial intelligence, which has made building simple bots accessible to far more people. So-called "basic" bots are gaining ground again, precisely because they've become easy to write.
The wave of AI bots
A recent phenomenon is shaking up the counters: the crawlers that feed artificial intelligence models. They comb the web to collect text, either for training or to answer a question posed in an assistant in real time.
Their growth is spectacular. The share of GPTBot, OpenAI's crawler, has more than doubled in a year within the verified-bot traffic measured by Cloudflare.
What's striking isn't just the volume, but the imbalance. Cloudflare published a telling ratio: the number of pages scraped by a crawler for each visitor it later sends back to the site. By the summer of 2025, that ratio was hitting peaks for some assistants, on the order of tens of thousands of pages crawled per single visitor referred. The company itself qualifies the measurement, because some applications don't pass along the information needed to count incoming visits, which mechanically inflates the ratio. The direction stays clear: sites give a lot, and get little back.
It's this imbalance that pushed Cloudflare, in July 2025, to block AI crawlers by default on the sites it protects, and to launch a system where access can be monetized. A shift from a web that's open by default to a web that asks permission.
Some sectors take a harder hit than others
The average hides wide disparities. Wherever there's value to extract or divert, bots concentrate.
- Online commerce: on retail sites, the share of bad bots reaches 59% of traffic according to Imperva. Competitor price monitoring, stock buyouts, stolen-card testing.
- Travel: roughly a quarter of all bot attacks target this sector, hungry for real-time fares and availability.
- APIs: 44% of advanced bot traffic targets programmatic interfaces directly, often less well protected than regular pages.
What it costs, in concrete terms
Behind the percentages, there are bills. Malicious bots weigh heavily on the online economy.
Let's add a few reference points, keeping in mind that some figures come from corporate communications and should be read as statements:
- A Netacea study estimates that bots cost large companies an average of 4.3% of online revenue.
- On single sign-on providers, one login attempt in five is "credential stuffing," the mass replay of stolen credentials (Verizon DBIR 2025).
- On the ticketing and sneakers side, the reported volumes are staggering: several billion bot attempts blocked each month by the largest players.
A word of caution before quoting these figures
What to take away from it
The web is no longer mostly human, and the automated share grows year after year. This reality has two opposite consequences. For those protecting a site, it justifies ever finer defenses. For those collecting public data legitimately, it explains why a simple script no longer cuts it: it now operates in an environment designed to tell, on every request, the machine from the person. Understanding that boundary is the subject of our other articles.
Collect this data without getting blocked
WyndPath handles proxies, JavaScript rendering and anti-bot bypass in a single API call. Pay-per-success.
Start for free →