News Classification Rules Being Developed for English and German with IPTC Media Topics
The IPTC has reached the first milestone in EXTRA, the Google/DNI project to build an open source rules engine for news. We are partnering with Infalia PC and have selected the Elasticsearch engine for developing a high-performance, rules-based news classifier. We are licensing an English language news corpus from Reuters and one in German from the Austrian Press Agency for use within the project. We have two linguists creating sample rules for classifying those corpora with IPTC’s Media Topics using the EXTRA engine. The project is on track to deliver a working version of the engine, together with the sample rules, by the summer of 2017.
EXTRA Open Source Rules for News
EXTRA (“EXTraction Rules Apparatus”) is an open source project to classify news text using rules. The engine allows news organizations to precisely identify the categories to which a piece of news belongs by specifying Boolean rules, with sophisticated natural language processing capabilities. Rule-based classification is better for breaking news than statistical methods, since it doesn’t require re-training using example news items (which typically take time to produce). Automated classification is generally more consistent and scalable than hand tagging of news. Most machine learning techniques are essentially “black boxes”, whereas rules provide much greater transparency – and therefore ability to control – why a piece of content is classified in a particular way. For all of these reasons, we believe that the EXTRA rules engine is ideally suited for news classification.
After evaluating a number of open source frameworks, we decided to make Elasticsearch’s percolator technology the foundation for the EXTRA engine. Our testing indicates that Elasticsearch supports indexing a large number of rules. The percolator has performant and scalable support for matching indexed rules against incoming documents, the core task of the EXTRA engine. Elasticsearch has an active open source community, as well as options for commercial support.
The EXTRA Requirements, Design, API and Rules Language
We have drawn up a detailed set of technical requirements and have created a high level technical architecture for EXTRA. We have designed the EXTRA API and the rule language. Linguists are working on writing the rules to classify English and German news using IPTC’s Media Topics taxonomy
IPTC, Infalia, Google DNI
EXTRA is being developed by the IPTC, an international consortium of news agencies, publishers and system vendors. The project is funded by the Digital News Initiative, Google’s €150 million fund aimed at stimulating innovation amongst European publishers. In 2016, IPTC applied for and won a DNI grant of €50,000 to develop the EXTRA engine. As a development partner, IPTC selected Infalia PC, a spin-out from the Information Technologies Institute of the Centre for Research and Technology Hellas with significant expertise in data analytics and natural language processing.
If you’d like to learn more about the IPTC or the EXTRA project, please contact firstname.lastname@example.org
IPTC’s Photo Metadata Working Group has released the Cultural Heritage Panel plugin for Adobe Bridge, which focuses on fields relevant for images of artwork and other physical objects, such as artifacts, historical monuments, and books and manuscripts.
Sarah Saunders and Greg Reser, experts from the cultural heritage sector, conceived the IPTC Cultural Heritage Panel to address needs of the photo business and growing community of museums, art foundations, libraries, and archive organisations. Furthermore the panel fills a gap: Many imaging software products, including Bridge, do not support all metadata fields of the IPTC Photo Metadata Standard 2016 for artwork or objects.
The artwork or object fields – a special set of metadata fields developed by IPTC a few years ago – describe artworks and objects portrayed in the image (for example, a painting by Leonardo da Vinci). This means that descriptive and rights information about artworks or objects is recorded separately from information about the digital image in which they are shown. Multiple layers of rights and attribution can be expressed – copyright in the photo may be owned by a photographer or museum, while the copyright in the painting is owned by an artist or estate.
The new plugin for Bridge (CC versions up to 2016 and CS6 were tested) allows people to view the image data, and write into these fields using a simple panel, which has been tailor-made for use in the heritage sector. The panel includes fields for artwork/object attributes and also relevant digital image rights.
“The Cultural Heritage Panel will be very useful for people working in the heritage sector in museums and archives,” Saunders, a consultant specialising in digital imaging and archiving. “It allows them to manage and monitor data about objects and artworks that is embedded in the IPTC XMP fields in the image.”
“The metadata can then be transferred into an organisation’s digital asset management system; the panel helps ease the ingest process,” Reser said.
Reser also noted that the panel helps incorporate more people into workflows, such as freelance photographers, who otherwise may not have access to an organisation’s digital asset management system. The Cultural Heritage Panel allows them to be an efficient part of the process of viewing the metadata included with an image, and adding to it when appropriate.
“IPTC is the most popular schema in embedded metadata,” Reser said. “Over time I bet we’ll see a lot of the cultural heritage fields creep into off-the-shelf programs and software.”
The panel is free, includes an easy-to-use interface, and includes key image administration fields. Image caption and keywords can be automatically generated from existing Artwork or Object data.
Download the IPTC Cultural Heritage Panel and User Guide for Adobe Bridge.
The IPTC has released a comprehensive set of sports controlled vocabularies as a supplement to the SportsML 3.0 sports-data interchange format, which was released in July 2016. These controlled vocabularies (CVs) are in the format of NewsML-G2 NewsML-G2 Knowledge Items plus RDF variants and are available on IPTC’s CV server at http://cv.iptc.org/newscodes.
There are 113 CVs representing such core sports concerns such as event and player status, as well as specialized lists for 11 sports (basketball, soccer, rugby, American football, etc.) for statistics, player positions, scoring types, etc.
“The SportsML 3.0 standard’s semantic tech capabilities are improved greatly by the new controlled vocabularies,” said Trond Husø, system developer for Norwegian news agency NTB, one of the early adopters of SportsML 3.0. “Data can be easily imported, structured, and stored.”
“When building a sports app you spend a lot of prep time defining your terms and building a schema,” said Paul Kelly, news technology consultant and lead for IPTC’s Sports Content Working Group. “By using SportsML 3.0, there is no need to reinvent the wheel.”
“You consider things such as ‘What sort of results and stats do we need?’ and ‘How will our system handle interrupted matches?’ IPTC’s vocabularies can get you on your way because they properly define in a standard format almost all the terminology you would use in a sports application: Everything from “goals-scored” to a full enumeration of status codes for sports events,” Kelly said.
For the Summer 2016 Olympics, NTB acquired the rights to distribute the results and data from the International Olympics Committee’s Olympic Data Feed (ODF). NTB then transformed ODF to SportsML 3.0, and then to NITF3.2. “Using SportsML to structure the ODF’s data is a broad and comprehensive solution to approaching all sports and competitions worldwide,” said Husø, who is also a member of IPTC’s Sports Content Working Group. “SportsML is now a truly flexible and universal format that can incorporate multiple vendor codes and still provide a defense against vendor lock-in.”
“Terms defined in another format such as ODF can easily live beside SportsML terms – as well as any other proprietary format – so that an organisation can build a repository of knowledge of all the different sports-data formats,” Kelly said.
Another advantage to the new SportsML 3.0 standard is that if new concepts are added to a sports vocabulary or modified in it, the data model and the XML Schema don’t change; they stay stable. It also supports all languages for the concept labels.
“A great feature is that we can translate the definitions to Norwegian – without changing or breaking the vocabulary,” said Husø. “If we were to distribute internationally, our domestic receivers could look up the definitions in Norwegian, while the international ones could use the English term.”
IPTC’s SportsML 3.0 standard underwent a major upgrade from version 2.2, after 12 years of evolution since its first version. The new standard incorporates contribution from sports experts in 12 countries. Its flexible core covers all major sports and events in most news reporting.
Other early adopters of SportsML 3.0 include Univision and the British Press Association in its new multi-sport API. Its major features include:
- compliance with IPTC’s NewsML-G2 standard
- a flexible core that covers all major sports and events in most news reporting
- plugins for detailed stats in 10+ sports
- a more flexible tournament model
- schedules, scores, standing, statistics, etc.
- choices between specific and generic terms
- controlled vocabularies, semantic tech capabilities
- schema redesign
- many samples and tool support.
Tool support for SportsML 3.0 includes 45 samples from 11 different sports and events, including both classic and SportsML-G2 examples, and both generic and specific examples.
The vocabularies will be maintained by IPTC for future expansion; new sports and terms can be added.
IPTC has published an updated Photo Metadata User Guide, for photographers, photo editors and professionals responsible for in-house metadata workflows, including digital asset management (DAM) systems.
Based on IPTC’s widely used Photo Metadata Standard, the new User Guide contains practical information regarding photo metadata – from photographers familiarizing themselves with basics, to managers in related businesses who have a deeper understanding of implementation of standards and metadata.
A key use of metadata is to describe the content of an image, location and rights information; the guide groups metadata fields according to information types. “The User Guide will help when deciding where metadata should be put about a certain topic, and what data should or should not be filled into a specific field,” said Michael Steidl, managing director of IPTC, and lead of IPTC’s Photo Metadata Working Group.
IPTC sets the industry standard for administrative, descriptive, and copyright information about images. The IPTC Photo Metadata Standard, supported by many software applications, is the most widely used standard because of its universal acceptance among photographers, distributors, news organisations, archivists, and developers.
The Photo Metadata User Guide walks users through the major groups of metadata, and for each IPTC field contained within each, it provides short guidelines on the use and semantics.
The first section of the guide outlines practical use for a basic understanding of applying photo metadata, and may be most helpful to photographers becoming familiar with adding it to their photos for the first time. Photo metadata is key to protecting photographers’ images, including copyright and licensing information online.
The User Guide addresses typical questions such as:
- What is a minimum set of fields to be used?
- How is metadata preserved?
Five examples of metadata for independent, staff, and agency photographers plus images of artwork are given.
Photo metadata is also essential for managing digital assets. Detailed and accurate descriptions about images ensure they can be easily and efficiently retrieved via search, by users or machine-readable code. This results in smoother workflow within organisations, more precise tracking of images, and potential for licensing opportunities.
For professionals responsible for in-house photo metadata workflows and DAM systems, all IPTC metadata fields in the User Guide have been grouped by topic for easy reference: general description, persons, locations, things shown, rights and licensing information, and administrative data.
The User Guide does not focus on the user interface of a specific software, and will be updated regularly to include more details.
IPTC is looking for software developers to design, develop, document and test EXTRA, an open source rules-based classification engine for news. First preference will be given to applications received by 21st October 2016, and review will continue until the positions are filled.
“Classification” means assigning one or more categories to the text of a news document. Rules-based classifiers use a set of Boolean rules, rather than machine-learning or statistical techniques, to determine which categories to apply.
EXTRA is the EXTraction Rules Apparatus, a multilingual open-source platform for rules-based classification of news content. IPTC was awarded a grant of €50,000 from the first round of Google’s Digital News Initiative Innovation Fund to build and freely distribute the initial version of EXTRA. DNI granted IPTC €50,000 for the entire project.
We are working with news providers to supply sets of news documents and with linguists to write rules to classify the documents. IPTC is looking for qualified developers to create the rules engine to accurately and efficiently categorize the documents using the rules.
Please consult this page for more information and to let us know if you’re interested in being considered.
The IPTC NewsCodes family of controlled vocabularies has a new member: Product Genre.
The Product Genre vocabulary was developed at the request of the broadcast industry. A broad category of terms was needed – one that specifies the kind of content by media product type – in addition to metadata that describes the content. The Product Genre scheme includes terms such as comedy, drama, entertainment, travel and sport.
NewsCodes are sets of concepts created and maintained by the IPTC, also known as controlled vocabulary or taxonomy. They are assigned as metadata values to news objects like text, photographs, graphics, audio and video files and streams. This allows for a consistent coding of news metadata across news providers and over the course of time.
The Product Genre vocabulary was an idea initiated by Andy Read, IPTC delegate and BBC’s Service Development and Delivery Manager for News, who has worked with IPTC for more than 20 years. This was based on feedback from broadcast members that highlighted the value of the forum engagements in driving the progression of the data set.
“There was a need to extend the breadth of the controlled vocabularies,” said Read. “The new Product Genre vocabulary codes describe the type of program itself, and help to broaden the program to a wider audience and general TV/broadcast industry.”
NewsCodes vocabularies can be very specific. A broader category like Product Genre allows identification of an entire broadcast program or package – not just smaller segments. For example, a longer 60-minute program overview about Syria’s war can be coded according to Product Genre – supplemented by metadata specific to a minute-long clip about a possible chemical attack, in the context of the larger news program.
“The Product Genre needed to be added to help facilitate use of these codes with IPTC’s NewsML-G2 standards,” said Read.
The new Product Genre vocabulary is also beneficial on the business side, said Jennifer Parrucci, senior taxonomist for the New York Times.
“Advertising is often sold based on the type of program – not necessarily subject tags or more specific terms,” Parrucci said. “The Product Genre vocabulary identifies advertising opportunities at a more comprehensive level.”
The IPTC NewsCodes Working Group, chaired by Parrucci, collaborated to define the vocabulary terms, based on concrete examples and actual TV programs. For each Concept identifier and name, a definition is listed. The notes section gives an example of what that Concept describes, for clarity and accurate use.
Any NewsCode provided by the IPTC can be used at any stage of a news workflow, without any royalty fee. But if one includes IPTC NewsCodes into an application, the intellectual property and the copyright of the IPTC must be explicitly attributed.
Interesting stats and info about the International Press Telecommunications Council’s technical standards for exchange of news information:
1.) The International Press Telecommunications Council publishes 14+ technical standards that are intended for the business-to-business exchange of news among news agencies, other news providers and publishers.
2.) At least one or two IPTC standards are in use at virtually every newspaper and news web site in the world.
Publishers use IPTC standards to save money and improve the ability of their news products to be used by customers.
3.) IPTC standards for news exchange are available for downloading at no cost – and there are no royalties or fees.
The only source of income for IPTC is membership dues. Membership currently consists of more than 50 organizations and individuals worldwide.
4.) All IPTC standards are designed to be independent of any specific language.
Although our publications are written in English and meetings are conducted in English, every recent standard is usable by any written language that is supported by Unicode.
5.) More than 70 software applications support IPTC Standards.
Software developers seamlessly integrate IPTC standards into their products – often in subtle ways that are not obvious to customers.
It’s an Olympic year for IPTC’s SportsML 3.0 standard, the recently released update to the most comprehensive tech-industry XML format for sports data.
“We figured, why not use the latest technology available?” said Trond Husø, system developer for NTB, who worked on the standard’s update, released in July. “SportsML 3.0’s use of controlled vocabularies for sport competitions and other subjects now provides many benefits, including more flexibility. Storing results is also more convenient.”
SportsML 3.0 is the ideal structure and back-end solution used by many major news organizations because it is the only open global standard for scores, schedules, standings and statistics. “It saves the time and cost of developing an in-house structure,” said Husø, also a member of IPTC’s Sports Content Working Party.
The Rio Games, which will host about 10,500 athletes from 206 countries, for 17 days and 306 events, are revolutionary for big data and new approaches for managing it. For the first time, the International Olympic Committee (IOC) used cloud-based solutions for work processes including volunteer recruitment and accreditation.
And consider the experimental technologies and apps launched by key broadcasters and Olympic Broadcasting Services, the Olympic committee responsible for coordinating TV coverage of the Games: virtual reality footage, online streaming, automated reporting, drone cameras, and Super-High Vision, which is supposedly 16 times clearer than HD.
Billions of Olympic spectators worldwide have naturally come to expect real-time results and accurate scores to be delivered to them, with a side of historical perspective. All with little thought as to how the information reaches the public, be it via tickers on websites, graphic stats on TV screens, or factoids offered by commentators.
Schedules, competitors’ names, bio information, times, rankings, medalists – how does all of this data get served up so quickly and uniformly among networks and news services? And how does it get integrated into existing news systems, namely SportsML 3.0?
It starts with the IOC – the non-profit, non-governmental body that organizes the Olympic Games and Youth Olympic Games. They act as a catalyst for collaboration for all parities involved, from athletes, organiser committees, and IT, to broadcast partners and United Nations agencies. The IOC generates revenue for the Olympic Movement through several major marketing efforts, including the sale of broadcast rights.
The IOC produces the Olympic Data Feed (ODF), the repository of live data about past and current games. The IOC is responsible for communicating the official results; they use the specific ODF format for their ODF data.
Paying media partners sign a licensing agreement to use ODF, to report on results through their own channels, and build new apps, services and analysis tools.
The goal of ODF is to define a unified set of messages valid for all sports and several different news systems – so that all partners are receiving the same data, at the same time. It was introduced for the Vancouver Games in 2010 and is an ongoing development effort.
According to the IOC’s website, ODF plays the part of messenger. From a technical standpoint, the data is machine-readable. ODF sends sports information from the moment it is generated to its final destination via Extensible Markup Language (XML). XML, a framework for storing metadata about files, is a flexible means to electronically share structured data via the Internet, as well as via corporate networks.
IPTC’s SportsML 3.0 easily imports data from ODF. Using SportsML to structure the ODF’s data is a broad and comprehensive solution to approaching all sports and competitions worldwide. ODF has identifiers for sports and awards (gold, silver, and bronze medals) executed at the Olympic Games; sports outside of ODF are identified by vocabulary terms of SportsML.
“SportsML 3.0 provides one structure for the data for developers to work in,” said Husø. “The structure will be the same, even if there are changes to ODF in future Olympic Games; the import and export process of the data will not change.”
Among content providers that use SportsML (various versions) are NTB, AP mobile (USA), BBC (UK), ESPN (USA), PA – Press Association (UK), Univision (USA, Mexico), Yahoo! Sports (USA), and Austria Presse Agentur (APA) (Austria), and XML Team Solutions (Canada).
SportsML 3.0 is based on its parent standard, NewsML-G2, the backbone of many news systems, and a single format for exchanging text, images, video, audio news and event or sports data – and packages thereof. SportsML 3.0 is fully compatibility with IPTC G2 structures.
Media Topics is an IPTC standard – a 1,100-term taxonomy with a focus on categorizing text. Released in 2010 as a development based on the IPTC Subject Codes, use of Media Topics is free and available in different formats. They can be viewed on the IPTC Controlled Vocabulary server, or in a user-friendly tree hierarchy tool.
IPTC creates and maintains taxomonies and controlled vocabularies – to assign terms as metadata values to news objects like text, photographs, graphics, audio and video files and streams. This allows for a consistent coding of news metadata across news providers, over the course of time.
“The idea of semantic mapping and being involved in a linked data initiative like Wikidata is a natural step for IPTC,” said Jennifer Parrucci, chair of the IPTC NewsCodes Working Group and senior taxonomist for The New York Times. “When linking an existing taxonomy to another, Wikidata serves as a central point of reference.”
Wikidata is a free, collaborative, multilingual knowledge base that can be read and edited by both humans and machines. It provides centralized storage for an access to structured data for all Wikimedia projects, as well as for use on external websites.
In total about 100 mappings from Media Topics to Wikidata have been manually applied. The mappings use SKOS mapping relationships.
Media Topics began with the Subject Codes vocabulary and extended the tree from 3 to 5 levels and reused the same 17 top-level terms. The lower-level terms have been revised and rearranged. Each Media Topic provides a mapping back to one of the Subject Codes.
The International Press Telecommunications Council (IPTC) is close to finalizing a new recommendation for video standards: the IPTC Video Metadata Hub.
The Video Metadata Working Group, which is comprised of members worldwide from news organisations, vendors and experts in the metadata field, is planning to vote on a recommendation of the Video Metadata Hub (VMD Hub) at the IPTC Autumn Meeting, 24 – 26 October 2016, in Berlin. The final Draft #4 has been published for a last round of reviews: http://dev.iptc.org/Video-Metadata.
Because there are several different existing standards for video – for compressing video and audio, file formats and different schemas of metadata properties – IPTC is presenting a “hub” recommendation that covers many use cases and exchange of metadata over multiple standards.
The VMD Hub is comprised of a single set of video metadata properties, which can be expressed by multiple technical standards (namely XMP for metadata embedded into binary video files, and EBU Core for non-embedded metadata stored in sidecar files). These properties can be used for describing the visible and audible content, rights data, administrative details and technical characteristics of a video.
Likewise, the VMD Hub supports workflow, exchange of metadata, and search functions across other existing standards, and will include mapping to Apple Quicktime, PBCore, MPEG7 and Schema.org, and perhaps more in the future.
“Users of videos of different standards told IPTC they need a common ground in metadata for efficient workflows,” said Michael Steidl, Managing Director of IPTC. “This is what we deliver now with the Video Metadata Hub.”
The IPTC Autumn Meeting will feature a Video Day on 25 October. In addition to the presentation about the VMD Hub, speakers from video makers, video suppliers, video content publishers and system vendors will discuss how video workflows can be improved.
For information about attending the IPTC Autumn Meeting and Video Day, contact us.