The global top 10 news sites for retaining embedded photo metadata, October 2026 edition.
The global top 10 news sites for retaining embedded photo metadata, October 2026 edition.
The IPTC has launched Metawatch, a free, open monitor that checks every month whether more than 500 news publishers in 122 countries keep the credits, captions and rights information embedded in their photographs.

The first results show how little has changed since 2018, which publishers have fixed their sites, and why keeping metadata matters more than ever now that AI companies are crawling the web.

Credits are lost between camera and reader

Almost every professional news photograph leaves the camera, or the agency, carrying information about itself: who took it, when and where, the caption, the credit line, the copyright and the licensing terms. Standards such as IPTC Photo Metadata and Exif exist so that this information travels inside the file, wherever the image goes.

In practice, most of it is gone by the time a reader sees the picture. Content management systems re-encode images, content delivery networks resize them, and optimisation tools discard anything that looks like unnecessary bytes. The photographer’s name and the agency’s rights statement are often the first things to go.

This has always been the case, but today it matters more than it used to. When an image is copied, shared or collected by an AI crawler, the embedded metadata is often the only record of where it came from and on what terms it may be used.

What Metawatch does

Metawatch checks every month: does the photo metadata survive publication? On the first of each month our well-behaved crawler visits 508 news publishers in 122 countries, takes up to 20 recent articles from each, and reads the lead photograph of every article. It checks each image for Exif, IPTC Photo Metadata (in both IIM and XMP form) and C2PA Content Credentials, and scores how each site carry the fields that matter most: creator, caption, copyright and credit line (the “4 Cs”).

The crawler is deliberately light: about 40 to 60 requests per publisher per month, honouring robots.txt and keeping no article text or image files. How it works is documented in full, and every run is published as open data under CC BY 4.0, so anyone can check our numbers or ask their own questions.

Metawatch builds on IMATAG’s 2018 State of Image Metadata study and IPTC’s own prototype crawler from 2019, presented at the 2019 Photo Metadata Conference. This service is different in that we now measure every month, so we can compare trends over time.

What we found

Eight years on from the IMATAG study, the main finding hasn’t changed. In 2018 IMATAG found that 8% of editorial images carried a credit or copyright notice. Applying the same test to our October run gives 7.2%. Whatever else has changed in news publishing since then, keeping the photographer’s credit attached to the photograph has not improved.

Most publishers keep nothing. In the October run, 321 of the 425 publishers we could crawl, three out of four kept none of the IPTC fields we score on any image we sampled. Only 11% of the 6,580 images carried any IPTC metadata at all. The field that survives most often is the photographer’s name, on just 7.9% of images; the credit line is on 5.7% and copyright on 5.6%.

It isn’t the CDN’s fault. Images served through Cloudflare, Akamai, CloudFront and Fastly lost their metadata 87% to 95% of the time, but images served with no CDN at all also lost metadata 87% of the time. Stripping appears to happen in the publishing pipeline, which means that publishers can fix it.

Some publishers have improved. USA Today carried no credit metadata when IMATAG measured it in 2019; in September 95% of its lead photographs carried credit, creator and copyright. Les Echos went from 6% to 75%. Going the other way, the Washington Post, HuffPost UK and El Pais carried credits in 2019 and showed none in our samples. Across the 294 publishers we have measured every month since June, the number keeping any metadata rose from 60 to 67. It is early, but the needle can move.

Which countries do best

France tops our October ranking, with an average publisher score of 19.8 out of 100, just ahead of Germany at 19.2. Taiwan enters at third. We rank only countries where we scored more than three publishers, so that one well-run newsroom cannot carry a whole country; 34 countries qualified this month.

Chart showing which countries are best at retaining embedded photo metadata in news sites. France is #1, Germany #2 and Taiwan is the biggest climber into the top 10 at #3, up from #15 last month.
Chart showing which countries are best at retaining embedded photo metadata in news sites. France is #1, Germany #2 and Taiwan is the biggest climber into the top 10 at #3, up from #15 last month.

The chart is why Metawatch runs monthly. The United States led in September and slipped to fourth in October, but only because two of its best-scoring publishers were out of reach: AP and USA Today began blocking automated access this month. It is a reminder that bot-blocking aimed at AI also shuts out research. Over time, the lines will show which countries’ newsrooms are changing their practice and which are standing still.

The IPTC Metawatch top 10 countries list for October 2026.
The IPTC Metawatch top 10 countries list for October 2026.
Even the leaders are a long way from good. A country average of 20 means most of its publishers still strip their photos; the high scores come from a handful of titles such as taz and Süddeutsche Zeitung in Germany, and Les Echos and Le Figaro in France.

AI: a locked front door, an open back door

News publishers are increasingly explicit about AI. Of the 505 publishers whose policies we checked, 234 (46%) block at least one AI crawler in their robots.txt. GPTBot, for example, is blocked by 36%.

But robots.txt protects a website, not a photograph. Once an image has been copied, syndicated or scraped from somewhere else, the only machine-readable statement of the publisher’s wishes is the one embedded in the file. IPTC Photo Metadata has a field for exactly this: PLUS Data Mining. Only 15 publishers (3.0%) put it on any image we sampled, and it appears on 0.6% of images overall.

Content Credentials (C2PA) are rarer still. Leaving out IPTC’s own site, 27 of the 6,580 images we analysed carried a C2PA manifest, and all but one were signed by an AI or design tool: OpenAI, Adobe, Canva or Google (the last had a signer we couldn’t read). None came from a camera or a newsroom’s own signing. Seven images carried an IPTC Digital Source Type saying they had been edited with generative AI, which is exactly the kind of disclosure the field was designed for. Provenance in news photography is arriving first through the tools that make synthetic images, not through the photojournalism that most needs it.

What publishers can do

The fix is usually a setting, not a project. Most image pipelines strip metadata by default to save a few kilobytes; most can be told to keep it. Start with the lead photograph and five fields: creator, credit line, caption, copyright notice and, increasingly, the data-mining preference. Then check the result on Metawatch the following month.

  • Publishers: find your site on Metawatch to see exactly which fields survive, image by image. If your site is showing as “blocked”, configure your bot blocker to allow “Metawatch” to see your content. If there is a problem with your site’s data, or if your site is missing from our list, tell us through the feedback form.
  • Researchers and journalists: every run is open data, free to reuse with credit.
  • Technology vendors: if your CMS, CDN or image service strips metadata by default, we can help you to fix it: please get in touch.

Metawatch automatically publishes new results on the first of every month. We hope to see the lines on the chart start to move upwards in the future.

Categories

Archives