Your Disease Surveillance Data Is Only as Useful as Its Fine Print
A new international guideline aims to fix a quiet problem in veterinary epidemiology: mountains of disease data that other researchers can barely use, because no one wrote down the details that would make it trustworthy.
Veterinary epidemiology runs on data — herd health records, outbreak logs, surveillance databases, laboratory results — collected across farms, clinics, and countries. In theory, all of that data should be a shared resource, reusable by the next researcher trying to model the next outbreak. In practice, a lot of it sits underused, not because it's bad data, but because nobody wrote down enough about how it was collected, cleaned, or defined for anyone else to trust it or reuse it correctly.
A newly published paper in Scientific Data, led by Céline Faverjon of EpiMundi in Lyon and co-authored by a working group of more than 20 epidemiologists and data scientists across Europe, Canada, and the United States, tries to close that gap. The result is the first community-built set of “rich metadata” guidelines built specifically for veterinary epidemiology — a practical framework, complete with templates and worked examples, for describing animal health data well enough that someone else can actually use it.
The Problem: Good Data, Missing Instructions
The FAIR principles — data that is Findable, Accessible, Interoperable, and Reusable — have been widely adopted across science for years. The trouble, the authors argue, is that generic metadata standards were never built with veterinary epidemiology's particular headaches in mind: multi-scale data that spans individual animals, herds, farms, regions, and countries, often pulled together from very different sources with very different quality standards.
In practice, the field has tended to put its documentation effort into describing the analysis and the models built on top of the data, while the raw data itself — where it came from, how complete it is, what its limitations are — gets far less attention. That imbalance is what quietly erodes reusability: a beautifully documented model built on a dataset nobody described properly is a model other researchers can't safely build on, verify, or extend.
What the New Guidelines Actually Do
Rather than inventing a new standard from scratch, the working group's approach was to integrate what already exists — established metadata schemes — with new, domain-specific guidance on the attributes and quality markers that actually matter for animal health data. The result is meant to function as a single-source framework: one place researchers across academic institutions, government agencies, and private industry can go to figure out what “well-documented” looks like for their dataset, rather than piecing it together from generic templates that don't quite fit.
Crucially, the guidelines didn't come out of a single institution's internal style guide — they were tested against real datasets by research teams across the collaborating institutions before publication, with contributors ranging from the Norwegian Veterinary Institute and the Swedish Veterinary Agency to Cornell, the Atlantic Veterinary College at the University of Prince Edward Island, Utrecht University, and the UN Food and Agriculture Organization. That range of contributors matters: a framework built and stress-tested across that many national surveillance systems and institutional data cultures has a much better shot at actually getting adopted than one built in isolation.
Why This Should Matter to Working Veterinarians
It's easy to read a metadata paper as something purely for data scientists and epidemiologists building the next disease model. But the practical stakes reach into clinical and field practice more directly than the title suggests. Every time a foreign animal disease response, a herd health surveillance program, or an outbreak investigation draws on shared data — from a regional lab network, a state veterinarian's office, a multi-clinic practice group — the speed and accuracy of that response depends on whether the underlying data can actually be trusted and combined with other sources. Poorly documented data slows everything down at exactly the moment when speed matters most.
There's also a longer-term payoff for anyone contributing data upstream — practices participating in surveillance networks, university diagnostic labs, industry researchers running herd health studies. Data described using a shared, rigorous standard is data that's far more likely to get reused in the next study, the next outbreak model, or the next policy decision, rather than quietly sitting unused because no one downstream could figure out what it actually meant.
An Open Framework, Built to Be Used
The paper is published open access under a Creative Commons Attribution license, meaning the guidelines, templates, and examples are freely available for any research group, lab, or surveillance program to adopt and adapt. The work was funded in part through European Union Horizon 2020 and European Partnership on Animal Health and Welfare programs, reflecting a broader push across the EU to standardize how animal disease data is collected and shared across borders — a goal that, if it takes hold, could make cross-border outbreak response meaningfully faster than it is today.
This account is based on “Enhancing reusability of veterinary epidemiological data by creating contextual metadata,” by Faverjon, Delavenne, Bareille et al., published in Scientific Data (2026), https://doi.org/10.1038/s41597-026-08217-9. Licensed under CC BY 4.0.
Share This Article
Free Membership
Enjoyed this article?
There's a lot more where that came from.
Join 50,000+ veterinary professionals who get free RACE-approved CE, weekly clinical updates, and the most talked-about veterinary magazine in the profession — all completely free.
Join Vet Candy Free →No credit card. No catch. Just everything veterinary.

