ANNUAL REPORT

Methodology

Customer total traffic analysis

Our primary dataset includes anonymized, de-identified traffic data across DataDome’s global customer base. For the 2026 report, we analyzed trillions of requests across 75,000+ customer sites between July 2025 and June 2026.

 

The data covers total traffic composition, attack vector volumes and trends, bot behavior patterns, and attack technique evolution. No end-user data is collected or stored. All figures represent aggregated, anonymized traffic signals.

How to interpret attack vectors data

Throughout this report, an attack vector refers to an automated request or sequence of requests that our systems classify as an attempt to carry out a specific activity, such as scraping, credential stuffing, distributed denial-of-service (DDoS), or fake account creation. It does not indicate that the activity was completed successfully or that a business impact occurred.

 

Attack vector classifications are based on signals observed in the request and session, including the endpoint targeted, request velocity and volume, behavioral patterns, browser and device fingerprints, network characteristics, claimed identity, and known attack signatures. A request may therefore be classified as an attempted attack even when DataDome blocks it before the action is completed. The figures in this section represent detected attempts, not confirmed successful attacks.

Customer AI traffic analysis

For our AI traffic analysis, we filtered our first dataset, looking specifically at AI traffic across DataDome’s global customer base. It classifies AI agents, agentic browsers, and LLM crawler traffic by bot type, endpoint, and referral source. The AI traffic covers 52.7 billion requests across DataDome’s 75,000+ customer websites. 

 

The AI endpoint and AI referrer analyses both cover H1 2026 (January–June 2026).

Test of 20k+ popular websites

For the third portion of the report, we created a dataset by testing 20,000+ of the world’s most popular websites across 15 industries and several major geographic regions (excluding those without a large enough sample size). Testing took place during June 2026. Domains were excluded if they were inaccessible during the testing window. The final valid sample was 21,491 domains, up from 16,948 in 2025. 

 

The study was conducted by testing websites’ homepages against multiple bot and AI traffic types by simulating real attack scenarios using residential proxies in three geographic locations: the United States, Canada, and France. Requests are routed through residential internet protocol (IP) addresses to avoid detection based purely on IP reputation.

 

The test sends requests from 10 bot types across four tiers of difficulty:

 

Tier 1: Basic bots (2 test types in tier)

These bots do not attempt to disguise themselves. They send raw, unadorned requests with minimal or generic headers. They are the lowest-effort attack type and the easiest to detect.

 

Tier 2: Spoofed AI agents and crawlers (4 test types in tier)

These bots forge the Hypertext Transfer Protocol (HTTP) headers of known, trusted identities, including Google’s crawler and AI agents from OpenAI and Anthropic. This includes spoofed versions of ChatGPT, GPTBot, and ClaudeBot. 

 

Tier 3: Disguised bots (1 test type in tier) 

The bot used in Tier 3 goes beyond headers and forges the network-level fingerprint of a real browser — including Transport Layer Security (TLS) handshake patterns and HTTP/2 behavior — without actually running one. Catching them requires looking past the visitor’s claimed identity and analyzing the connection itself against known browser fingerprint profiles.

 

Tier 4: Real browsers (3 test types in tier) 

These bots run a genuine browser engine, such as Chrome or Firefox, and are fully automated and equipped with anti-detection techniques designed to mimic human behavior. These are the most sophisticated bots and the hardest to detect.

 

Scoring framework: Each domain receives a score based on how many bot types were challenged or blocked. A site is scored fully protected if it challenged or blocked all tested bot types, partially protected if it challenged or blocked some — but not all — test requests, and unprotected if it let all 10 test bot requests through without a challenge. 

 

Fully protected: Successfully blocked or challenged 100% of our ten bot types.

Partially protected: Successfully blocked or challenged a portion of our test bots.

Unprotected: Failed to block or challenge any test bot.

A note on the ethical use of bots

Our test is designed to have zero impact on sites. Tests are rate-limited, non-destructive, and designed to simulate realistic bot and AI behavior, not to stress or disrupt infrastructure. No end-user data is accessed or retained. The test does not exploit vulnerabilities; it only tests whether automated traffic triggers a detection response.