From Fragmented Scientific Data to an AI-Ready Research Ecosyste

Chatgpt Image Aug 13, 2026, 03 26 40 Pm

Summary

BioTeam partnered with an NIH Institute to assess the FAIR-ness (Findability, Accessibility, Interoperability, and Reusability) and long-term sustainability of its genomics data and software resource portfolio. BioTeam designed and deployed a FAIR self-assessment questionnaire, conducted structured interviews with resource-holding Principal Investigators, synthesized findings across dozens of federally funded resources, and delivered a tiered set of recommendations to close FAIR gaps and strengthen ecosystem-wide interoperability.

The Institute needed an evidence-based picture of how well its funded data and knowledge resources were actually performing against FAIR principles, beyond individual resource-level assessments or claims. The resulting assessment established a quantified, portfolio-wide baseline, identified metadata, identifier, ontology, and API practices that were contributing to fragmentation, and produced a prioritized roadmap for building a more coherent, sustainable, and AI-ready research data ecosystem.

Challenge

The Institute funds a large and diverse portfolio of data repositories, knowledge bases, software resources, and data coordinating centers. These resources had been developed independently over many years by different teams, for different scientific communities, with different technical priorities and approaches to data stewardship.

There was no consistent way to measure or compare how FAIR individual resources actually were, nor a shared understanding of where the most significant interoperability and sustainability gaps existed across the broader ecosystem.

Specific challenges included:

  • Inconsistent metadata practices. Resources handling related genomic and clinical data used different metadata standards, with no shared data model connecting them across the portfolio.
  • Fragmented identifier strategies. Resources were divided between established persistent identifiers such as DOIs and internally generated identifiers, without consistent guidance about which digital entities should receive persistent identifiers.
  • Uneven ontology and vocabulary practices. Investment in ontologies and controlled vocabularies varied substantially, including differences in staffing and resources dedicated to their development and maintenance.
  • No consistent API strategy. APIs were viewed by nearly all participating resources as essential to their function, but there was no shared standard or strategy governing their development and adoption across the ecosystem.
  • Long-term sustainability concerns. Resources were developed and maintained independently, with limited coordination around long-term data stewardship, archival planning, and sustainability.

The Institute needed to move beyond anecdotal, resource-by-resource impressions of FAIR-ness and establish an evidence-based, ecosystem-wide assessment that could inform future priorities, funding decisions, and policy.

Approach

BioTeam designed the engagement as a two-phase assessment combining quantitative self-assessment data with qualitative insights from resource leaders.

Phase 1: FAIR Self-Assessment at Scale

BioTeam designed and deployed a FAIR self-assessment questionnaire to dozens of Institute-funded resources spanning data repositories, knowledge bases, software resources, and consortia-level data coordinating centers.

The questionnaire captured FAIR-specific measures while also examining broader sustainability characteristics drawn from federally recognized best practices for scientific data repositories. This allowed BioTeam to evaluate not only how resources aligned with the four FAIR principles, but also the organizational and technical practices that could affect their long-term sustainability and usefulness.

Phase 2: Structured Principal Investigator Interviews

Using the questionnaire results, BioTeam selected a representative cross-section of Principal Investigators for structured, in-depth interviews. Participants represented different data types, levels of informatics maturity, resource models, and approaches to data access.

The interviews provided context that quantitative responses alone could not capture. BioTeam explored how resource leaders approached metadata, persistent identifiers, ontologies, APIs, data stewardship, and sustainability, as well as how they were thinking about emerging requirements for AI-ready scientific data.

The interviews also surfaced areas where existing FAIR frameworks did not map cleanly to particular resource types, providing important context for interpreting the assessment results.

Synthesis, Analysis, and Recommendations

BioTeam combined the quantitative questionnaire results with qualitative findings from the interviews to create a comprehensive view of the Institute’s research data ecosystem.

Findings were organized around the four FAIR pillars and examined both resource-level practices and recurring patterns across the broader portfolio. BioTeam then developed a tiered set of major and supplemental recommendations, prioritizing actions based on their potential impact and feasibility.

Rather than treating FAIR as a checklist for individual resources, the assessment identified opportunities for the Institute to improve FAIR-ness and interoperability at the ecosystem level.

Outcomes

The assessment gave the Institute an objective, portfolio-wide baseline of FAIR-ness, replacing resource-by-resource impressions with a structured and comparable view across dozens of federally funded scientific resources.

The analysis identified several recurring sources of ecosystem fragmentation, including meaningful inconsistencies in persistent identifier practices, substantial variation in the staffing and investment supporting ontologies and controlled vocabularies, and near-universal reliance on APIs without a shared approach to their development or adoption.

Together, these findings helped explain why individually valuable scientific resources could remain difficult to connect, discover, and use together across the broader research ecosystem.

BioTeam translated those findings into a prioritized path forward for Institute leadership, including recommendations to:

  • Explore shared metadata models for resources managing related scientific data
  • Establish clearer guidance around persistent identifier selection and use
  • Encourage greater standardization and adoption of APIs across resources
  • Strengthen coordinated ontology and controlled-vocabulary practices
  • Improve long-term approaches to data stewardship and sustainability
  • Establish ongoing FAIR self-assessment practices to measure progress over time

The engagement moved the Institute from a fragmented, resource-by-resource understanding of its scientific data portfolio toward a coordinated, evidence-based strategy for improving FAIR-ness, interoperability, and long-term sustainability.

Just as importantly, it identified the metadata, identifier, ontology, and API foundations needed to make scientific resources more discoverable, interoperable, and usable by emerging AI-enabled research workflows.

Share:

Newsletter

BioTeam updates, delivered.

Have Questions?

We'd love to help.