Can AI help authorities identify hazardous chemicals earlier?
NEWS
New research from Umeå University shows that AI tools can convert chemical structures in patent images into searchable digital data with up to 78 per cent accuracy, a promising step towards early warning systems that flag potentially hazardous chemicals before they reach the market. But the technology still has clear limitations on complex structures and poor-quality images.
Patents are far more than legal documents protecting intellectual property, they are a vast source of scientific information about chemicals that are under development, often years before they appear in products or in the environment. That makes patents a valuable data source for Early Warning Systems (EWS), which authorities use to identify potentially hazardous chemicals before they become a threat to human health and the environment.
There is just one problem: much of the chemical information in patents is locked inside drawings. Molecular structures and reaction schemes are embedded as images that no text search can read. Less than 7 per cent of the patents examined in the study contained chemical structure images at all, but when they do, the structures are essentially invisible to computers unless they are converted into machine-readable formats.
– A patent might depict a new molecule years before it is produced at scale," says Farina Tariq, doctoral researcher at the Department of Chemistry, Umeå University. – If we can teach computers to read those drawings, we can give scientists and authorities an earlier opportunity to investigate potentially concerning chemicals. But first we need to know how reliable the AI tools actually are when applied to real patent documents, not just clean images from databases.
The study received additional recognition when artwork based on Farina Tariq’s scientific poster was selected for the cover of the journal issue. The cover highlights both the research and the visual communication of its central idea: using AI to extract chemical information from patent images.
Image[Simon Jönsson]
From picture to searchable data
The potential workflow is straightforward: a patent is published → an AI system detects a chemical structure in an image → the structure is converted into a machine-readable format → the system compares it with known chemicals and chemical classes → potentially concerning or entirely novel structures are flagged for expert investigation.
Once a structure is machine-readable, it can be checked against known hazards: Is this chemical already known? Is it structurally related to a known hazardous substance? Does it contain structural features associated with particular hazards? Or is it an entirely new structure for which little hazard information exists? If so, can we compute its potential hazards?
Putting three AI tools to the test
In the study, published in the journal Chemical Research in Toxicology, the research team tested three widely used chemical structure recognition tools, on two specially curated data sets from the European Patent Office's Espacenet database: one of general organic chemistry and one of per- and polyfluoroalkyl substances, PFAS. The AI-generated structures were then validated by five chemistry experts.
The results were two-sided. On standard, well-drawn organic structures, all three tools performed well, with accuracy around 74–78 per cent, correctly decoding aromatic rings, common functional groups and standard abbreviations. On the PFAS data set, however, performance dropped sharply: for 26 of the 43 unique PFAS structures, not a single tool produced a correct result.
While the AI tools could accurately interpret many chemical structures, they had difficulties with some of the complex ways chemicals are described in patents. Older patent documents, where images may be blurred or distorted, also led to mistakes.
– If an AI system misreads a molecular structure, everything that follows is built on faulty information, says Farina Tariq. – An incorrect structure could trigger unnecessary risk mitigation measures — or worse, let a truly concerning chemical slip through unnoticed. That's why quality control and confidence scoring must be built into any future automated workflow.
The researchers stress that the goal is not for AI to declare a chemical dangerous. Rather, AI acts as a translator: converting chemical drawings into digital structures that can be compared with known chemicals, screened for worrying structural features, and prioritised for expert evaluation.
Image[Simon Jönsson]
– Our conclusion is not that the technology is ready to be switched on tomorrow, says Patrik Andersson, professor at the Department of Chemistry and co-author of the study. – But we now have a realistic picture of where these tools succeed and where they fail. That is exactly what is needed before automated screening of patents can strengthen early warning of hazardous chemicals in the future.