Skip to content

Ad Libraries Are the Only Social Data a Regulator Made Public

Every other public dataset on a platform is public by permission, and permission can be withdrawn. Ad transparency archives exist because the law requires them.

11 min read · 20 Aug 2026

If you are choosing what social data to build a product on, almost every option carries the same unpriced risk: the data is public because a platform decided it should be, and that decision can be reversed on a Tuesday with no notice, no migration path, and no appeal. Nitter, the Reddit API repricing, Instagram's logged-out endpoints, X's API tiers — the pattern is the same each time. You did not lose access because you did something wrong. You lost it because the permission you were relying on was never a commitment.

Ad transparency archives are the exception, and the exception is structural rather than lucky. Meta, TikTok, Google, LinkedIn, X, Snapchat and Pinterest publish searchable ad repositories because Article 39 of the EU Digital Services Act obliges them to, on pain of fines calculated as a percentage of global turnover. That is a different kind of public, and it changes the risk arithmetic for anyone deciding where to point their crawler.

This post is about that difference — what the mandate actually says, which risks it removes, which it emphatically does not remove, and the ways the archives are worse than the pitch suggests.

This is not legal advice. It is a developer's reading of a regulation and some platform documentation, written by engineers. What is lawful depends on your jurisdiction, your data, and what you do with it. Talk to a lawyer about your specific situation.

On the feature question: if what you want is a field-by-field comparison of the seven ad libraries — which exposes creative, which exposes reach, which retains for how long — adlibrary.com's seven-platform comparison and their DSA developer guide already do that thoroughly and we are not going to write a worse version. They are a commercial ad-data vendor, so read the recommendations accordingly, but the platform tables are good. This post covers the thing those guides deliberately leave out, which is risk.

What Article 39 actually requires

The DSA's transparency obligation applies to very large online platforms and search engines — Article 33 sets the threshold at 45 million average monthly active recipients in the EU, with designation by the Commission rather than self-assessment. Once designated, Article 39 requires a provider that shows ads to:

"compile and make publicly available in a specific section of their online interface, through a searchable and reliable tool that allows multicriteria queries and through application programming interfaces, a repository…"

Four load-bearing details in that sentence and the paragraphs that follow it:

Requirement What it means for you
"through application programming interfaces" Programmatic access is part of the obligation, not a favour
"searchable… multicriteria queries" Filtering by advertiser, date, country and format is required, not a feature
Retained for the display period plus one year There is a floor under how fast data disappears
Contents: ad creative, on whose behalf, who paid if different, dates, targeting parameters, and total reach broken down by Member State The field list is a legal minimum, not a product decision

Article 39(1) adds a constraint that matters more than it looks: the repository must be compiled without the personal data of the recipients of the ad. The archives were designed from the start to exclude the people who saw the ad, which is not true of any other social dataset you might scrape.

And Article 74 supplies the reason any of it is honoured: fines of up to 6% of total worldwide annual turnover of the undertaking — the group, not the local entity. That is the enforcement mechanism standing behind the availability of this data. It is a considerably stronger guarantee than a platform's changelog.

"Public by mandate" versus "public by accident"

Here is the distinction that should change what you build.

Public by accident describes almost everything else. A profile page loads without a login because the platform wants it indexed, wants the growth loop, or has not got around to closing it. The data is reachable, so people reach it. But nothing obliges the platform to keep serving it, and the entire legal analysis — the CFAA, the terms of service, the case law everyone quotes — is about what happens when the platform decides it would rather you didn't.

Public by mandate inverts every one of those questions:

  • There is no authorisation question, because you are an intended reader. Article 39 exists so that researchers, journalists and the public can query this data. Reading it is the purpose of the thing.
  • There is no technical measure to get past, which matters more in 2026 than it did in 2022. The newest litigation theory in this area is not the CFAA but circumvention — evading rate limits and anti-bot measures. Where the platform is required to publish an API, there is nothing to circumvent.
  • The platform cannot quietly withdraw it. Removing the archive is a compliance failure with a percentage-of-turnover price tag, not a product decision. Compare that to any undocumented endpoint you currently depend on.
  • The data does not concern the audience. Article 39(1) keeps recipient personal data out by construction, which removes the hardest category of GDPR exposure before you write a line of code.

That last point is worth being precise about, because it is where the argument is usually oversold.

Where the GDPR question gets easier, and where it doesn't

Easier: the repository contains ads and advertisers, not the people who were shown them. The single most awkward conversation in social data — why are you holding a database of identifiable individuals who never interacted with you? — mostly does not arise. There is no follower list, no comment author, no profile of a private person.

Not gone: an advertiser can be a natural person. A byline naming who paid for a political ad is, in EU terms, personal data about that person. Targeting parameters are not recipient data, but they describe how a population was segmented. And processing personal data lawfully requires a lawful basis regardless of how the data reached you — a point that survives every US ruling in this area, and the one the whole category most reliably elides.

The honest summary is that ad-library data is the easiest GDPR conversation available in social data, not the absence of one.

The trade you make by using the official API

Almost nobody says this part out loud, so here it is.

The argument for logged-out scraping is that you never accepted anything — no account, no click-through terms, no contract to breach. Using an official ad-library API gives that up. You register a developer app, you agree to platform terms, and you now have a contractual relationship with the platform you are reading.

That is a real trade and we think it is a good one, for a specific reason: the terms you accept are published, versioned, and attached to a documented interface with a deprecation policy, rather than being an unwritten expectation you might be violating without knowing. You swap "no contract, unknown exposure" for "a contract you can read, plus a stable interface." But it is a swap, not a free win, and anyone telling you the official API is risk-free is skipping a step. Read the platform terms before you build, particularly the clauses about resale and redistribution — most of these APIs have them, and they are the constraint that actually bites commercial products.

Which archives you can reach programmatically

Compressed, because the feature detail is covered better elsewhere and it changes often. What matters for a build decision is the access model, not the field list.

Archive Programmatic access Gate
Meta (Facebook, Instagram, Threads, Messenger) Official Graph API (ads_archive) Identity confirmation, a developer app with ads_read
TikTok Commercial Content API Formal application; approved researchers, non-commercial commitment
Google Ads Transparency Center; a BigQuery dataset for political ads Browse UI, or the public dataset
LinkedIn None Browser only — no official API, no bulk export
X, Snapchat, Pinterest DSA repositories, browse-first Little to nothing for commercial ads outside the EU

Meta is the only one with a mature, generally available API for commercial ads, which is why it is the one most products end up on and why it gets its own field-by-field reference.

The pattern underneath the table is worth noticing: five of these seven archives are compliance artefacts and behave like it. They exist because Article 39 required them, they cover EU-delivered ads properly, and outside the EU they publish close to nothing about commercial advertising. The DSA created an asymmetry where an analyst in Berlin can see targeting parameters and per-country reach for a campaign that an analyst in New York, researching the same advertiser, cannot see at all.

The EEA-only problem, if your market is the US

This is the single biggest practical disappointment for US-focused products, and it follows directly from the fact that a European regulation is doing the work.

The richest fields in Meta's archive fall into two groups. One group — reach breakdowns by age, country and gender, the beneficiary and payer of the ad, total EU reach — is documented as available only for ads delivered in the EU or UK. The other group — spend ranges, impression ranges, demographic distribution, estimated audience size, funding bylines — is available only for ads about social issues, elections or politics.

If you are building a US commercial ad-intelligence product, you get neither. You get creative text, link titles and captions, delivery start and stop dates, publisher platforms, languages, the page that ran it, and a snapshot URL. That is enough to answer "what is this advertiser running and since when." It is not enough to answer "how much are they spending," and no amount of engineering will make it so, because the data was never collected into the archive in the first place.

Design around this early. A roadmap that assumes spend data will arrive once you get the token is a roadmap built on a field that does not exist for your ads.

The regulator giveth, and the regulator taketh away

The strongest argument for building on ad-library data is that a regulator guarantees its existence. The honest counterweight is that a regulator also controls its scope, and in 2025 it narrowed it dramatically.

The EU's Transparency and Targeting of Political Advertising regulation (TTPA) applied from October 2025. Rather than comply, both major platforms exited the category:

  • Google stopped serving political advertising in the EU, with YouTube ceasing on 23 September 2025.
  • Meta announced the same in 2024 and confirmed it on 6 October 2025: "As of today, political, electoral and social issue adverts are no longer able to be delivered in the EU." Meta cited the TTPA's "significant operational challenges and legal uncertainties."

Follow that through. The fields with real analytical value — spend, impressions, demographic distribution — exist only for political and issue ads. The EU has the strongest transparency mandate in the world. And as of October 2025 there are no new political ads being served in the EU to be transparent about. The archive keeps its history, and the flow stopped.

The general lesson is not "the DSA failed." It is that a mandate guarantees publication of what is served, not that anything will be served. A dataset created by regulation can be shaped by regulation, including out of existence, and your product plan should treat the field set as a policy variable rather than a constant. That is still a far better position than depending on an undocumented endpoint, but it is not the risk-free foundation it is sometimes sold as.

What no ad library will ever tell you

Worth stating plainly, because a surprising number of products get built on the assumption that it is in there somewhere:

  • Performance. No clicks, no CTR, no conversions, no ROAS. Impression and spend ranges exist for political ads and are wide enough to be nearly useless for competitive analysis — you learn that a campaign spent somewhere in a bracket, not what it earned.
  • Exact spend, for anything. Ranges only, political only.
  • Creative files. Meta gives you a snapshot URL that renders the ad; it does not hand you the image or video assets.
  • The call-to-action button. Not a field. If you need it, you are reading a rendered page, which puts you straight back into the fragile category.
  • Landing-page content, funnel structure, or anything after the click.

An ad library tells you what an advertiser said and where they said it. It tells you nothing about whether it worked. Every "we reverse-engineered their winning ad" claim built on this data is inference from how long an ad ran, which is a real signal and a weak one.

If you are going to build on one thing in this category, build on this

Set against the alternatives, the case is straightforward. Ad-library data has a legal guarantee of continued publication, an official documented interface, no anti-bot surface, no proxy requirement, no authorisation question, and no recipient personal data. Nothing else in social data has more than two of those.

It also has genuine commercial pull. Ad-intelligence tools charge real money for access to this data — AdSpy at $149/month and Foreplay from $49/month as of August 2026 — which is a reasonable proxy for what the underlying archive is worth to someone.

Our own position, stated so you can discount it appropriately: we read the Meta Ad Library through Meta's official Graph API, from our own IP, with no proxies and effectively zero marginal cost. It is the highest-margin data in our product for exactly the reasons above. And as of this writing our four Ad Library endpoints are not live — they are built and blocked on an access token, and they return a clean not_configured error naming the missing variable rather than pretending otherwise. We would rather say that than imply a working integration we do not yet have.

The failure modes we are watching, and you should too: the token expires and has to be regenerated; identity confirmation takes days and can be refused; the field set is a policy variable, as October 2025 demonstrated; and the whole approach depends on continuing to be an intended reader rather than an unwelcome one. That is a shorter and much less exciting risk register than the one attached to any other social platform, which is the entire point.

The mechanical half of the argument — why an HTML scraper pointed at the Ad Library UI breaks repeatedly while the Graph API keeps working — is covered separately, along with the cases where scraping genuinely still wins.

Everything described here runs on the same API. You are never charged for a failed request, an empty result, or a cache hit.