This summer has taken its toll on farmers across the world. Improving resilience and adapting fast to this ever changing environment will depend on building more resilient soils. This will require better evidence about what works, where it works and under which conditions. Yet much of the soil data needed to build that evidence remains difficult to find and access.
Large volumes of soil data are generated every year by governments, research institutions, development programmes and private organisations. But these datasets are often held in separate systems, described using different terminology and stored in incompatible formats. Some are publicly available; others are subject to legitimate restrictions concerning ownership, privacy, licensing or national data sovereignty.
A conventional response would be to bring all this information into one central repository. For soil data, however, centralisation is neither always practical nor always desirable. Data owners may be unable or unwilling to transfer their data to a third party. Regulations and governance arrangements vary between countries and institutions, and moving data into a central platform can mean losing control over how it is accessed and used.
A better approach is to make independently managed systems interoperable. Through a federated architecture, organisations can operate their own infrastructure and retain control of their data while publishing standardised metadata that makes relevant datasets discoverable across a wider network. Access to the underlying data can then remain open, restricted or subject to approval according to the data owner’s own rules.
It’s a model that’s gaining traction across industries. But it’s not without its challenges.
The engineering challenge of interoperability without centralisation
The first challenge is helping independently managed systems to understand one another. Soil data may be stored in different formats, described using different terminology and collected according to different methodologies. Shared standards for data, metadata and APIs are therefore essential, but the data must also be harmonised carefully enough to remain scientifically meaningful.
Those standards will adapt as scientific understanding, technology and user needs change. A federated system must therefore be able to evolve while maintaining compatibility between different versions. Otherwise, every change could force participating organisations to update their systems at the same time, creating a significant barrier to adoption.
Access control also becomes more complex when there is no single organisation managing the entire network. Each participating node remains responsible for protecting its data and deciding who may access it. Solving this requires more than a technical architecture. It also requires agreement on standards, governance and the responsibilities of every organisation participating in the network.
What building critical infrastructure in the open changes for software teams
A federated architecture addresses how independently managed systems can work together. But there is also a question of trust: why should governments, research institutions and other data custodians build their systems around technology controlled exclusively by one organisation?
Open-source software increases trust by making all the details of the implementation known. Participating organisations can inspect the code, deploy it within their own infrastructure and gain full understanding of how it works. Equally they are not required to hand control of their data or technical operations to a single platform provider. This makes openness particularly important for infrastructure intended to work across different countries and institutions.
Opening the code also changes the responsibilities of the original software team. Public scrutiny can help uncover defects and security vulnerabilities, but it must be supported by a disciplined process for reporting issues and developing fixes and integrating contributions.
Its role therefore expands beyond writing code. The original team must help build and sustain a community by responding to requests, supporting contributors, documenting decisions and managing discussions about the platform’s direction. External contributions do not remove the need for a core team; they create a broader set of relationships and responsibilities for that team to manage.
Open source creates the conditions for community-driven collaboration, but publishing the code alone is not enough. A healthy community depends on active maintenance, clear contribution processes and open discussion of design choices.
The future of data infrastructure
This reflects a principle in software engineering known as Conway’s Law, where systems tend to mirror the structures of the organisations that create them. Within a company, that can mean software reflecting the boundaries between teams. At the scale of a global data ecosystem, the architecture must instead reflect the boundaries between independent institutions, regulatory environments and areas of responsibility.
In soil data, no single institution owns all the information or has authority to determine how every dataset should be used. A federated architecture is therefore not merely a technical choice; it reflects how ownership, trust and decision-making are distributed across the ecosystem.
Roberto Prato
Roberto Prato is Head of Engineering at Varda Foundation, a non-profit technology platform enabling farm and field data sharing to support the transition to a global nature-positive food system.


