Chapter Five: Agricultural Data Systems
Chapter Message
Agricultural sovereignty begins not with the model that predicts, but with the record that can answer: What was measured? Where and when? In what unit and with what instrument? Who collected it? What changed before it reached the screen? If those answers are lost, the observation becomes an orphaned number: easy to transfer and link, but poor in meaning and difficult to trust. Nor do data acquire value merely by accumulating. A raw value that appears wrong may be the only evidence of a sensor fault, a conversion error, or an exceptional event in the field. A mature system does not erase that trace in the name of cleaning; it preserves the original, distinguishes what was collected from what was normalised and standardised and from what was approved for use, and records every transformation, who or what performed it, and why. The result therefore remains open to review, reconstruction, and withdrawal if an error comes to light. This chapter therefore treats the data system as a structure of knowledge and rights together, not as a neutral repository of numbers. Farmers need to know how data connected to their farms have been used, for what purpose, and who can access them. They also need a practical way to download, transfer, correct, and leave the system without losing their agricultural history or becoming captive to a single vendor. When the record is traceable, the raw data are preserved, and the terms of use and exit are clear, artificial intelligence becomes an accountable tool; without these safeguards, it may multiply both the speed of error and the depth of vendor lock-in.
1. What Do We Mean by Agricultural Data?
Let us imagine a single season on a single farm. It begins with field boundaries drawn on a map, a soil analysis from a specified depth, and a rain forecast updated hourly. Satellite images and mobile-phone photographs then arrive, along with moisture readings from scattered sensors, pump operating hours, fertiliser and spray applications, and notes from a worker who saw yellowing at the end of one row. At harvest, yield, quality grades, and losses are added, followed by records of storage, transport, and price. These scattered traces describe the same field, but they view it from different angles and do not speak the same temporal or spatial language. All of these are agricultural data, but the category extends beyond plant images and soil measurements. It also includes stocks, costs, prices, wages, the movement of inputs and products, machinery, maintenance and energy records, animal health, cold-chain temperatures, laboratory results, observations by advisers and scouts, and local knowledge describing a growth stage, a weather sign, or an inherited practice. The data may take the form of a number in a table, text, an image, coordinates, a time series, an audio recording, or an event emitted by a machine operating in the field. Strictly speaking, agricultural information is not simply a value. It is a representation of something that happened, was described, measured, or carried out, linked to an entity, a time, a place, and a method. The number “20” does not tell us whether it means twenty degrees Celsius, twenty millimetres of rain, or a dose of twenty kilograms per hectare. A leaf image alone does not tell us which cultivar was photographed, at what growth stage, under what lighting, or whether it was taken before or after treatment. What appears to be an administrative detail about the data may in fact be the difference between a valid comparison and a misleading conclusion. Agricultural data can be viewed as overlapping families: data describing resources and the environment, such as land, soil, water, and weather; data monitoring biological conditions, such as plant growth, disease, and animal health; data recording action, such as irrigation, fertilisation, spraying, and harvesting; data describing outcomes, such as yield, quality, and loss; and economic and social data concerning price, cost, labour, ownership, and access. This division aids understanding, but it does not remove the connections among them. A single irrigation decision may combine the weather forecast, soil moisture, crop growth stage, energy price, the available share of water, and the worker's ability to act in time. These sources differ in scale, frequency, accuracy, and the consequences of error. A satellite image may summarise a wide area every few days, whereas a sensor reports from a single point every few minutes. A laboratory report describes a sample collected at a particular time, place, and depth, while a worker's observation remains linked to their experience, language, and the circumstances in which they saw the sign. Bringing these sources together in one file does not make them comparable; each may be sound within its own limits and lose meaning when detached from them. It is therefore not enough for a platform to say that it ‘integrates data’. It must first know the entity to which each observation belongs: a farm, a field, a management zone within a field, a plant, an animal, or a crop lot. It must preserve the observation's time, place, unit, and measurement method; the device or laboratory that produced it; calibration status; the identity of the person who created or entered it; and the version of the data dictionary or schema used to interpret it. It should also record the basis authorising its collection and use, the associated licence, the purpose for which it was collected, and whether that purpose permits reuse, sharing, or model training. These contexts are not decoration around the number; they are part of its meaning. Without them, a model may learn a relationship that does not exist in the field, compare two seasons recorded in different units, turn a missing measurement into zero, or attribute local knowledge to a system that did not produce it. When the original is preserved and identity, context, and rights are established, records can be connected without erasing their differences, turning a heap of files into agricultural knowledge that can be used and reviewed. This leads to the next question: what journey does an observation take from the moment it is collected until it becomes an approved record, enters a model, or is turned into a recommendation and a decision?
2. A Lifecycle, Not a Black Pipeline
When a recommendation appears on screen saying, ‘Add eighteen millimetres of water’, it can seem as though the result arrived fully formed. Yet this number may be the final link in a journey that began with a sensor reading, passed through a communications network, underwent a unit conversion, was linked to a field, crop, and growth stage, and entered a model before becoming a recommendation. If that journey cannot be traced backwards, the user cannot know whether the result came from a reliable measurement, an imputed missing value, a misinterpreted unit, or an old record wrongly attributed to the current season. A data system should therefore not be seen as a pipe into which values enter at one end and answers emerge at the other. An observation does not pass through the system exactly as it arrived; it may be transferred, copied, reformatted, converted to another unit, linked to an entity, combined with other sources, and used to derive a new variable. At each step its meaning may be enriched or changed, part of it may be lost, or an error may be introduced that was not present at the outset. The problem with a ‘black-box pipeline’ is not automation itself, but the inability to see what the automation has done. If the system changes a value, drops a row, imputes a missing element, or chooses between two possible interpretations, it must leave a trace showing what changed, when, under which rule, by whom or which service, and with what level of confidence. A transformation that leaves no record may make the data tidier, but it makes the resulting knowledge less verifiable. The lifecycle of an agricultural observation can be traced through four principal stages:
- Creation and collection: Where did the observation originate? Did it come from a sensor, laboratory, image, paper form, manual entry, or external system? Which device, method, or person created it?
- Transfer and preservation: How did the observation move from its source to the system? Did it arrive intact? Did its encoding, format, or timezone change? Was an exact copy of what arrived preserved, together with a checksum demonstrating its integrity?
- Transformation and linking: Which cleaning, standardisation, and verification operations were performed? Into which unit was it converted, and under what rule? To which field, crop, season, or entity was it linked? Was that link confirmed, probable, or in need of review?
- Use, publication, and disposition: Was the observation used in a report, an indicator, model training, or an operational recommendation? Who saw the result? How long is it retained? What happens to it and to products derived from it when it is corrected or withdrawn, when its purpose ends, or when deletion is requested?
Not all data need to pass through every stage in the same way. A corrupt reading may stop at the verification gate; a low-confidence observation may be retained for review without entering a model; and data may be approved for a research question yet remain unsuitable for an operational decision. What matters is that these states are not conflated, and that a record's arrival in the database does not become an implicit certificate of correctness or fitness for every purpose. To prevent this confusion, it is useful to separate three clear layers:
- Raw layer: The value is preserved as received, together with its original file or message, collection context, and checksum. Calling it ‘raw’ does not mean it is correct; it may contain an entry error, an anomalous reading, or an unknown unit. Its value lies in preserving the original trace to which reviewers can return, so it is not silently replaced even when it appears wrong.
- Normalised layer: This contains the value after its format, unit, or terminology has been standardised, with an explicit link to the original. If a temperature is converted from Fahrenheit to Celsius, a decimal separator is interpreted, or a local crop name is linked to a standard identifier, the transformation rule and version, input, result, and confidence must be recorded. Normalisation here is a documented interpretation, not a new fact that erases what came before.
- Approved layer: This contains records that have passed known verification and review rules and are fit for a defined purpose. Approval does not mean that a record is correct in every context; it may be suitable for a seasonal report but insufficient to operate a pump automatically, or suitable for aggregated statistical analysis but not for diagnosing an individual plant. Approval is therefore tied to purpose, risk level, the identity of the approving body, and the review date.
This separation of layers is not intended to multiply copies unnecessarily, but to prevent correction from erasing evidence. If an impossible value appears, the past is not rewritten to make the record look clean; the raw value is preserved, a corrected or normalised value is created, the reason for the correction is recorded, and a decision is then made on whether the observation merits approval, quarantine, or a request for a new measurement. Suppose that a rule for converting temperatures from Fahrenheit to Celsius was applied incorrectly to data from several greenhouses. In a black-box system, the fault may not become apparent until implausible alerts or ventilation commands are issued, and the team may be unable to identify which records and decisions were affected. In a traceable system, the defective rule version can be identified, every value that passed through it can be listed, the reports, models, and recommendations that relied on those values can be found, and the calculation can be rerun from the raw values without collecting the data again or guessing what they were before correction. The same applies to less obvious errors: converting tonnes per feddan into tonnes per hectare, interpreting the date 01/02/2026 as the first of February rather than the second of January, associating a village name with a similarly named field, treating an empty cell as zero, or combining soil measurements taken at different depths. Each change may look small in a table, yet its effects may extend to a quality indicator, a seasonal comparison, a forecasting model, or a decision with financial consequences. Derived products should therefore not retain only their final value. An average, indicator, forecast, recommendation, and even an AI summary should each retain links to the inputs on which it was built, the version of the rule or model used, the time of creation, and the confidence level. The provenance record then resembles a family tree: a reviewer can move from the result to its constituent elements, from a normalised element to its original, and from the original to the file, device, or person that created it. Preserving raw data does not mean keeping everything forever. Retention is itself a decision governed by need, risk, contract, law, and the rights of data subjects. Policy may require sensitive personal or commercial data to be deleted, minimised, or anonymised once the purpose has ended. Controlled deletion, however, differs from silent disappearance: the system must know what was deleted, under whose authority and when, and which derived products must be withdrawn or rebuilt, without retaining content that may no longer lawfully be kept. A trustworthy system does not claim that errors never enter it; it knows where an error entered, what it affected, and how it can be reversed. When the observation's journey from collection through use to archiving or deletion is visible, the data pipeline is no longer a black box; it becomes a chain that can be inspected, questioned, and reconstructed. We can then ask the more specific question: what elements must an observation record contain if its scientific, operational, and rights-related meaning is to remain intact?
3. An Example of a Complete, Meaningful, and Accountable Observation Record
If the data lifecycle reveals the route an observation has travelled, the observation record is its passport within the system. It preserves not only the value, but also what is needed to understand and test it, link it to its source, and determine what may be done with it. Let's imagine that a sensor installed in a tomato field sent a reading that the volumetric water content of the soil was 24%. The information seems clear at first glance, but it alone is not sufficient to answer the question that concerns the farmer: Does the field need irrigation now? Before using this reading, we need to know the position and depth of the sensor, the time of measurement, how the value was arrived at, the condition of the device, when it was last calibrated, and whether the reading is for the entire field or a small point of it. We also need to know what the system did with the original number, and whether the reading arrived on time, or remained stored in the device and then sent after the network was interrupted. The number 24% may be mathematically correct, but it remains meaningless unless we know whether it represents soil moisture in the root zone, air humidity, or a surface reading affected by recent irrigation. Knowing the moisture value alone does not mean irrigation is required; its interpretation depends on soil type, crop, growth stage, rooting depth, weather, irrigation method, and the moisture thresholds shown to be suitable for that field. The following table shows how a single value is transformed into an understandable and auditable agricultural observation. The example is educational and does not represent a specific farm or device.
| What does the record preserve? | Simplified example | Why is it needed? |
|---|---|---|
| Observation identifier | Reading No. 184 | Gives every reading a stable identifier so that it is not confused with another during correction or review. |
| What was measured? | Soil moisture | Explains what the number means. The value 24% might refer to soil moisture, air humidity, or a loss rate, each with a different meaning. |
| Field or area measured | Tomato field No. 7, eastern zone | Identifies the agricultural location described by the reading. A reading from one zone must not be generalised to the entire field without verification. |
| Sensor location | Near the third irrigation line, in the middle of the eastern zone | Helps the farmer or technician find the measurement point and compare it with surrounding plants. |
| Measurement depth | 20 centimetres below the soil surface | Surface moisture may differ from moisture in the root zone, and measurements from different depths must not be combined as though they described the same condition. |
| Measurement time | 12 August 2026, 7:30 a.m. farm time | Establishes when the reading was taken. A morning reading may differ from an afternoon reading, and yesterday's reading does not necessarily describe the field today. |
| Time received | Received by the system at 7:34 a.m. | Reveals transmission delay. If the network fails and a reading arrives hours later, the system must not present it as real-time. |
| The value as sent by the device | 0,24 | The original number is saved as it arrived from the device or file, before any modification or conversion. |
| Meaning of the value for the reader | Volumetric soil moisture of 24% | Expresses the number in a form that is easier to understand and compare, rather than leaving it as an opaque decimal. |
| Interpretation of the ratio | About 24 litres of water in every 100 litres of total soil volume | Explains the ratio in more intuitive language, without treating it alone as an instruction to irrigate. |
| Unit of measurement | Cubic metres of water per cubic metre of soil, displayed to the user as a percentage | It is forbidden to compare numbers that use different units, or to interpret an abstract number without knowing what it represents. |
| Measurement method | Moisture sensor installed inside the soil | It differentiates between a device reading, a laboratory result, a visual estimate, and a calculation derived from other data. |
| Device used | Moisture sensor No. 17 | Enables the device's readings to be identified if it is later found to have malfunctioned or produced biased values. |
| Calibration status | The device was checked on July 1, 2026, and its reading was within the acceptable range | It helps to estimate the extent of confidence in the measurement. The device may continue to send regular numbers, even though they deviate from reality due to poor calibration. |
| Reading source | Sent by sensor No. 17 through the farm gateway | Shows the route by which the information arrived, rather than attributing it vaguely to ‘the system’. |
| Location in the original file | August readings file, ‘Field 7’ sheet, row 148 | Enables a return to the original source so the reading as received can be reviewed if a discrepancy or error appears. |
| What did the system change? | Interpreted the comma in 0,24 as a decimal separator, then displayed the value as 24% | Makes the transformation visible. If the number-interpretation rule is wrong, the error can be found and the value recalculated from the original. |
| Reading status | Technically sound, but represents only one point in the field | Prevents a reading that is valid at its location from becoming a general judgement about a larger area. |
| Confidence | High for successful receipt and device calibration | States the dimension of confidence; it does not mean that the reading alone is sufficient for an irrigation decision. |
| Permitted use | Daily monitoring and support for review of the irrigation decision | Defines what may be done with the reading and prevents its automatic use for a higher-risk purpose for which it was not reviewed. |
| Prohibited use | Do not operate the pump automatically on this reading alone | Establishes a practical boundary preventing information from passing directly into an action that could harm the crop. |
| The holder of the right to the data | The farm, according to agreement with the system provider | Explains who has the authority to allow access, transfer, correction, or deletion of data. |
| Participation and training | Do not send to an outside party, nor use to train a public model without consent | It is prohibited to transfer the data to a new use that has not been approved by the right holder. |
| Retention period | It is stored for three years, then the need for it is reviewed | It is prohibited to retain data for an indefinite period, and its retention is linked to a clear purpose and policy. |
| Review status | The irrigation official reviewed it and approved it for the daily report only | It clarifies who has authorized its use, and in what scope, rather than considering it approved for all purposes. |
This example reveals the difference among three things that may look identical on screen: the value, the observation, and the approved record. The value is 0,24. The observation is that value linked to soil moisture, the tomato field, a depth of twenty centimetres, and half past seven in the morning. The approved record is the observation once its source, transformation, quality, and rights of use are known, and an authorised person has permitted its use for a defined purpose. Completing these fields does not mean that the observation is necessarily correct. The sensor may be known, the unit clear, and the location recorded, yet calibration may be wrong or the reading may fall outside the device's reliable range. The record remains complete in this case, but receives a quality status such as ‘Suspect’ or ‘Needs review’. This is better than deleting the reading: deletion conceals the problem, whereas the quality status preserves the evidence and prevents silent use. Likewise, the degree of trust should not turn into an absolute seal of health. The system may be confident that the column represents soil moisture, but it does not know that the sensor has moved from its position. The connection of the reading to the field may be correct, while the depth of measurement remains unknown. So confidence is most useful when you clarify the question it answers: Are we confident about the type of the property? Or from loneliness? Or from the field identity? Or from the safety of the device? Collecting all these questions into one number may hide the weakness rather than reveal it. It is also important to differentiate between the source of the observation and the right holder. The sensor generated the reading, the platform may have stored it, and the company may have developed the analysis tool, but that does not automatically give any of them the right to sell the data or use it to train another model. The record should state who has authorization authority, for what purpose it is permitted, for what period of retention, and what should happen when the contract expires or a transfer or deletion is requested. Note: How should the farmer view this record? The platform should not display all previous farm details at once. These fields are necessary in the system's background for tracking and review, but not all of them are suitable for the daily interface. The observation summary can be presented to the farmer in this form: Soil moisture: 24% Location: Tomato field No. 7, eastern sector Depth: 20 cm within the root zone Measurement time: 7:30 AM Device status: Checked, and the reading is technically correct Reading boundaries: Represents a single point, and does not describe the entire field Suggested action: Compare with other readings, plant condition, and rain forecast before making the watering decision In this view, farmers first see what they need: the value, location, time, reading status, and limits of use. Growers, technicians, or auditors who need more can open ‘Measurement Details’ to see the device, calibration, raw value, transformations, source, review, and rights of use. Good simplification neither removes evidence nor hides complexity that affects the decision; it arranges the information in layers. A clear, comprehensible summary appears at the surface, while enough detail remains behind it for verification, accountability, and reconstruction. The farmer is not forced to read a long technical record, and the specialist is not forced to trust a number whose origin cannot be verified. In this example, the platform does not tell the farmer, ‘Irrigate because moisture is 24%.’ It says, ‘This is a reading of 24%, taken at this place, depth, and time, with this device and at this confidence level. It is fit to support review of an irrigation decision, not to make that decision on its own.’ Data support agricultural judgement, but do not replace the context that gives them meaning.
4. Missing Is Not Zero
A farmer may open a weather report and find that yesterday's recorded rainfall was zero millimetres. They may open another report and find the same field left blank. The two states look alike on screen, but mean very different things: zero means a measurement was taken and no rain was detected, whereas a blank may mean that the weather station did not measure, the connection failed, the file did not arrive, the value was withheld, or the system did not know how to interpret it. This is not an accounting detail. If the system turns an empty cell into zero, it may conclude that the field received no rain and recommend irrigation even though rain fell while the station was out of service. If it turns a missing yield value into zero, the season may be recorded as a total failure although the crop was harvested and the quantity was simply not entered. If an unreported pesticide dose is treated as a zero dose, the system may conclude that the pest appeared in an untreated plot when treatment actually occurred but was not documented. Zero, then, is a meaningful value. Missingness describes the state of our knowledge about a value. Confusing the two does not merely fill a gap in a table; it changes the story the data tell about the field. States that look empty but do not mean the same thing A single mark, such as a dash or a blank cell, cannot express every kind of absence. At a minimum, a data system should distinguish the following states:
| State | Simplified agricultural example | What does it mean? | How should the system handle it? |
|---|---|---|---|
| Measured zero | The weather station measured rainfall and recorded zero millimetres | A measurement was made and the phenomenon did not occur, or its measured value was zero | Preserve 0 as a valid measurement, with its time and device |
| Not measured | There is no station at the site, or the worker did not take a reading | No observation exists | Record ‘Not measured’; do not replace it with zero |
| Measured, but below the detection limit | The laboratory tested a sample and found no pesticide residue above the detection limit | A measurement was made, but the value lies below what the method can detect | Record ‘Below the detection limit’ and the detection limit; do not treat it as a confirmed zero |
| Device failure | The moisture sensor stopped because its battery failed | A measurement was expected, but the instrument produced no valid reading | The fault state is recorded and a maintenance or alternative-measurement alert issued |
| Reading arrived late | The network failed and the device later sent the morning readings in the evening | The value exists, but was unavailable when the decision was made | Preserve measurement and receipt times; do not present the old reading as real-time |
| Invalid reading | A temperature sensor reported 89°C in an open field on a mild morning | A value exists, but validation rules indicate that it may be wrong | Preserve the raw value and mark it ‘Suspect’ or ‘Not approved for use’ |
| Not applicable | Pesticide-dose field for a plot not yet planted | The question does not apply to this case | Record ‘N/A’ so the field is not interpreted as a zero-dose treatment |
| Not yet entered | Harvest is complete, but the recorder has not entered the production quantity | The event occurred, but its data have not yet been recorded | Record ‘Awaiting entry’, together with the responsible person and due date |
| Unknown | A crop quantity is missing, and it cannot be established whether it was weighed at all | There is insufficient information to determine why the value is absent | Record ‘Unknown’ and do not assume a cause unsupported by evidence |
| Withheld value | The farm did not permit the selling price to be shared with the body preparing the report | The value may exist, but access is not permitted | Record ‘Withheld’, the reason, and access permissions without exposing the value |
| Deleted under policy | The retention period expired for personal data linked to a worker | The value existed and was deleted to enforce a right or policy | Record the fact, reason, and date of deletion without retaining the deleted content |
| Lost in transfer | The original file contains the reading, but it did not reach the database | The absence results from an import or connectivity error | Repeat the transfer or import and check whether other records were affected |
This distinction allows the system to say what it knows and what it does not know. When the phrase “not measured” appears, the farmer knows to measure. When it says “Device Failure,” the technician knows the problem is with the tool or connection. When “below the detection limit” appears, the researcher knows that the measurement has been completed, but the device cannot determine a more accurate value. If it appears “blocked”, the issue relates to the right of access, not the quality of measurement. How should the status appear to the farmer? Farmers do not need obscure technical codes or blank spaces that demand guesswork. The interface can display these states in direct language: Zero: The measurement was done and the result was zero. No reading: The measurement was not performed or its result was not received. below the detection limit: A measurement has been made, but the quantity is smaller than the instrument can determine. Not applicable: This measurement or procedure is not specific to this case. The device is stopped: The measurement could not be performed due to a malfunction or interruption. Suspect reading: A value arrived but did not pass the validation rules. Blocked: Data exists, but the current user is not authorized to see it. Colour or a symbol may help the reader, but colour should not be the only means of conveying meaning. A clear word matters more than a visual signal, especially in print, for someone with impaired vision, or on a small screen in the field. Imputation is not a neutral cleaning operation When values are missing, a system may try to replace them so it can complete an analysis. It may use the average of nearby readings, the last known value, a measurement from a nearby station, or an estimate from a statistical model. This is called ‘imputation’, and it is more than tidying a table: it creates a value that was not measured directly. It is therefore a modelling decision that must be visible and reviewable. If the temperatures recorded during a given day are 22, 23, a missing value, and 25 degrees, putting the average in the missing box may seem like a reasonable solution. But the missing value may have coincided with an hour when there was a short heat wave that raised the temperature to 35 degrees. In this case the average does not complete the record, but rather erases the most important event in it. Using the last known reading may be appropriate when a property changes slowly, but becomes misleading in a rapidly changing property. Soil moisture after irrigation begins, the temperature of the cooling room when the cooling device malfunctions, and the concentration of a gas inside a greenhouse can change within a short time. Repeating the previous value in these cases creates an imaginary stability that did not occur in reality. Even imputation by an advanced model does not turn an estimated value into a real measurement. The model may use weather, soil, and neighbouring readings to produce an estimate better than a simple average, but it remains an estimate based on assumptions. The imputed value must therefore remain clearly distinguishable from the measured value. If a missing value is replaced, system should log:
- That the original value was missing, and that the absence state should not be replaced with silence.
- Known reason for the absence, such as equipment failure, file delay, or failure to perform a measurement.
- The imputation method used, such as an average, the nearest station, or an estimation model.
- The data that was used to construct the estimated value.
- The version of the imputation rule or model that produced the value.
- The resulting value and the degree of uncertainty associated with it.
- The person or service that authorized its use.
- The purpose for which it is permitted to be used.
- Possibility of recalculation if the true value appears later.
Thus the system maintains two facts at the same time: There was no original reading, and A temporary estimate was generated according to a known method. If the blank is replaced with an unmarked estimated value, the user and the model will later treat it as a real measurement, and the boundaries between what was seen in the field and what the system inferred are lost. Sometimes not compensating is the safest decision Not every gap should be filled. If the information will be used to draw a general, low-risk trend, documented imputation may be acceptable. If it will operate a pump, adjust greenhouse ventilation, initiate a chemical treatment, or support a decision concerning animal health or food safety, the safest course may be to stop the automated decision and request a new measurement. Depending on the consequence of the error, the system can choose one of two clear paths:
- Follow up the analysis by showing that some values are estimated.
- Reducing the degree of confidence in the result.
- Request a new measurement or field verification.
- Use an independent alternative source.
- Submit a recommendation for human review without implementing it.
- Switch to a pre-defined safe operating mode.
- Stop the decision if the basic data is insufficient.
The right question is not: “How do we fill in each empty box?”, but rather: “Can this decision be made safely given what we do not know?” Absence itself may contain information Missing values are not always distributed by chance. Gaps may occur in locations with poor connectivity, on farms that cannot afford to maintain equipment, or during high-pressure seasons when workers are too busy to enter records. Some measurements may be lost because access to the field is difficult after rain, or because a device malfunctions at high temperatures. In these cases, the absence does not merely reflect a technical problem; rather, it may reflect disparities in infrastructure, resources, and recording capacity. If a model is trained on available data alone, it may learn better from farms and regions that are more connected and organized, and then perform poorly for areas that are already underserved. The algorithm may learn from the pattern of absence itself. If the disconnection is linked to a remote region, or the lack of production records is linked to small holdings, the algorithm may use the absence of data as an indirect reference to location or economic status. This does not mean that the absence pattern should always be deleted; it may be useful for anticipating network failures or routing support. But it does mean that its use must be conscious, and its effects examined, before it becomes a hidden source of bias. So missing value is asked at two levels. The first level is technical: Did the device crash, the network was down, or the import failed? The second level is social and operational: Are there regions, categories or seasons where gaps occur? Who does not appear in the data because collecting it is more difficult, more expensive, or inaccessible to them? Practical rule The rule can be summarized in four phrases: Zero is a measurement, not a space. Empty space is a condition that needs to be explained, not a number that needs to be invented. An estimated value is still an estimate, even if it is produced by the best model. When the error consequence is high, requesting a new measurement may be better than filling in the box. A good system is not ashamed to say, “We don’t know.” Rather, it shows whether the value was not measured, not arrived at, not discovered, not applied, withheld, or rejected. The limits of knowledge are part of knowledge itself, and a truthfully described void is more valuable than an accurate-looking zero that did not occur in the field. However, even when a value is present and not missing, its meaning may still be ambiguous due to different units, number formats, dates, and languages. This leads to the next challenge: How do we unify the data without corrupting its original meaning?
5. Units, Dates, and Languages
A farmer may send a file stating that the field covers ‘five feddans’, while another system records area in hectares and a third uses dunams. A laboratory may write a result as 1,250, which one program interprets as one and a quarter and another as one thousand two hundred and fifty. A date may appear as 01/02/2026 without revealing whether it means the first of February or the second of January. These problems may look merely formal, no more than differences in notation, yet they can change both the meaning of the data and the decision based on them. If field area is misinterpreted, calculations of seed, fertiliser, or water requirements will be wrong. If day and month are reversed, a treatment may be recorded as occurring before rather than after the onset of disease. If a decimal separator is misunderstood, a small dose may become a figure hundreds of times larger. Standardising agricultural data is therefore not limited to changing text formats or translating words. It is an interpretive process that must preserve the original meaning, document the transformation rule, and expose uncertainty rather than hide it. The unit does not come attached to the number A measurement has no complete meaning without its unit. The value 20 might mean twenty millimetres of rain, twenty litres of water per tree, twenty kilograms of fertiliser per hectare, or a temperature of twenty degrees Celsius. The number is mathematically valid in each case, but it leads to a different decision. Even familiar units may carry several meanings. A dunam does not represent the same area in every country or historical context. The definition of a quintal may differ from one country to another or from one crop to another. A tonne may mean a metric tonne or another unit used in some markets. A percentage in a soil analysis may also be calculated on a dry- or wet-weight basis; the % sign alone does not reveal which is intended. This is why the system must remember four separate things:
- The number is as stated in the source.
- The unit as written by the data subject.
- The meaning or definition associated with this unit.
- Standardized value after conversion, with the conversion rule used.
If a farmer writes that the field covers five feddans, the system does not silently replace the phrase with a value in hectares. It retains ‘5 feddans’ as the original value, records the country or the definition of the feddan used, and then calculates the standard area under a known rule. If the unit definition changes, or the country is found to have been specified incorrectly, the calculation can be rerun from the original. Conversions should be visible to the user in simple language, such as: Area as entered by the user: 5 feddans Normalised area: Approximately 2.10 hectares Conversion rule: The feddan used in this record equals 4,200 square metres Conversion status: Confirmed after specifying the country and source unit If the word ‘feddan’ or ‘dunam’ appears without the country or intended definition being known, the system should not choose a default conversion and conceal the assumption. The safe course is to preserve the original value, mark it ‘Unit requires clarification’, and request user review. Date is not a neutral number A similar problem appears with dates. The phrase 01/02/2026 may mean the first of February in the day/month/year system, or it may mean the second of January in the month/day/year system. If the date is related to agriculture, irrigation, or the appearance of a disease, this difference may reverse the order of events. History is best presented in words when it is directed to humans: February 1, 2026, 7:30 a.m. farm time Within the system, the date is preserved in a standardised format together with the timezone, calendar, and interpretation of the original text. ‘Seven o'clock’ is insufficient when data arrive from different timezones, and measurement time is not always the same as the time at which the value reaches the system. This becomes still more important when connectivity is interrupted. A sensor may measure soil moisture at seven in the morning, store the reading internally, and transmit it at two in the afternoon when the network returns. If the system records only the receipt time, the morning reading will appear to describe the soil's afternoon condition. A distinction must therefore be made between:
- Measurement-event time: The moment at which the device read the field condition.
- Recording time: The moment the device wrote the reading into its memory.
- Receipt time: The moment at which the reading reached the platform.
- Processing time: The moment at which the system verifies or converts the reading.
There are also agricultural dates that do not represent a single day. The phrase “2025/2026 season” may describe a production cycle extending over two years. The phrase “two weeks after germination” links time to an agricultural event, not to a fixed calendar date. The “flowering stage” is a vital state that may begin at different times between fields and varieties. Therefore the system should not force all agricultural times into a single date slot; rather, it preserves the difference between date, period, season, age after planting, and phenological stage. A comma may change the field value The way numbers are written varies between languages, countries and programs. The comma may be used to separate the integer part from the decimal part, or it may be used to separate the thousands. The point may perform the opposite function. The following table shows some examples of confusion:
| What is stated in the source | Possible meanings | Safe handling |
|---|---|---|
| 1,25 | One and a quarter, or a malformed value depending on the locale setting | Check the source country and locale before conversion |
| 1.250 | One and a quarter, or one thousand two hundred and fifty | Do not approve the number until the source convention for thousands and decimal separators is known |
| 1,250.50 | One thousand two hundred and fifty and a half in some systems | The original text is preserved and the reading rules are determined by the source |
| 1.250,50 | One thousand two hundred and fifty and a half in other systems | Automatic conversion is prohibited if the file language or region is unknown |
| 05/06/2026 | 5 June or 6 May | Require the day-month order to be stated or infer it only from other verified data |
| 24 without a unit | Temperature, moisture, dose, or another quantity | The value is quarantined until the property and unit are known |
| 5 dunams | An area whose correct conversion depends on the country or the adopted definition | Establish the local context before converting to square metres or hectares |
It is not enough for the system to guess the most common meaning. The guess may succeed in most rows, but then spoil the most important rows. So it preserves safe normalisation:
- The original text as stated.
- Language or locale likely.
- Reading selected by system.
- The rule used in the conversion.
- Degree of confidence in this interpretation.
- Any potential alternatives have not been ruled out.
- The result of human review when there is ambiguity.
If the reading is clear, it can be standardized automatically. If it has two reasonable meanings, the process should stop upon review, rather than the platform silently choosing one of them. Language carries knowledge, not just names Linguistic diversity is not confined to translating interface buttons or table headings. The names of crops, cultivars, diseases, pests, growth stages, and agricultural practices carry local histories and experience accumulated in a particular region. A farmer may use a common name for a symptom, an agricultural adviser a technical term, a laboratory the name of a pathogen, and a database a stable scientific identifier. These expressions are not always synonymous. A common name may describe a cluster of symptoms rather than a single confirmed disease, and a word's meaning may change between regions. A cultivar name may have a local pronunciation or several spellings, while the name registered with the competent authority remains stable. If a system replaces every local expression with a single standard term, its dictionary may look orderly while erasing distinctions needed for diagnosis or advisory work. This is why a unified reference entity should be built for each concept, carrying a fixed identifier that does not depend on the language or the way it is written. Then link to this entity:
- Scientific or official names.
- Common and local names.
- Different writing methods.
- Abbreviations.
- Language and dialect.
- The region in which the name is used.
- The intended definition in that context.
- The source of the term and who introduced it.
- Degree of confidence in matching.
- Whether conformity is confirmed or needs to be reviewed by a specialist.
A system might, for example, record a disease name as a farmer pronounces it and link it to a broader concept describing the symptoms, without automatically turning it into a confirmed diagnosis. The original wording remains preserved because what the farmer said is an observation, whereas identifying the disease is an inference that requires further evidence.
The same applies to class names. A local name is not deleted simply because a trade or registered name exists, and a trade name is not automatically treated as a distinct genetic variety. The system must know whether the name refers to a crop, a variety, a trademark, or a local designation for a group of varieties. Translation Does Not Mean Replacement When the platform translates an agricultural term, it must preserve the original word, the translation, and the context in which it was used. Translation may bring the meaning closer, but it may narrow a broad concept or expand a specific concept. In some cases it is better to display the local term alongside the reference name, rather than hiding it. The term can appear to the user like this: Name entered by the farmer: The local name used in the village Related reference concept: Symptoms of plant wilt Conformity status: Requires field verification Note: Matching does not represent a definitive diagnosis of a specific disease Thus, the system preserves the user's language, and at the same time makes use of the reference dictionary for searching and linking. Local knowledge does not become noise to be cleaned up, nor does a standard term become a tool for erasing difference. The direction of writing is part of the integrity of the presentation When Arabic, Persian, Turkish, or English symbols and numbers meet, the order of the line segments on the screen may change. The unit symbol may appear before or after the number in a confusing manner, parts of the date may be reversed, the device ID may be split, or a negative sign may be mixed up with the value. So the system needs to test type direction, especially in fields that combine Arabic text with Latin units, numbers, and identifiers. The meaning should not depend on a visual position that may change between the screen, the exported file, and printing. It is preferable to display the value, unit, and date in a clear format, keeping the technical identifiers separate from the readable description. What does the user see, and what does the system save? The farmer does not need to see all the coding details and conversion rules every time. The interface can display: Field area: 5 feddans, or approximately 2.10 hectares Measurement date: February 1, 2026 Soil moisture: 24% Term used locally: Saved in history Data status: Consolidated and reviewed In the background, the system maintains the original value, local unit, language, region, conversion rule, execution date, confidence score, and the identity of who reviewed the result. Thus, the user gets a simple presentation, without the specialist losing the ability to review the details. Practical rule Safe uniformity can be summarized in four rules: Do not convert a number before you know its unit and the context in which it is written. Do not interpret an ambiguous date based on the shape of the numbers alone. Do not replace a local name with a standard term; link them and preserve both. Do not let the system hide the guess: if the meaning is ambiguous, ask for a review. Good standardization not only makes all data similar on the surface, but also makes its differences understandable and relatable. It preserves what the source wrote, adds a documented interpretation to it, and distinguishes what is confirmed from what is likely. Then two systems can exchange data without exchanging errors with it. However, agreement on the date format or unit symbol does not guarantee that the two systems mean the same thing. Both may use the word “region,” while one means an administrative region and the other means a sector within a field.
6. Semantic Interoperability
Two systems may exchange a file without any technical error and still fail to exchange meaning. The second system opens the CSV file, reads its rows and columns, and recognises its numbers and dates, so the transfer appears successful. Yet a column called ‘region’ may mean an administrative region in the first system, an agricultural field in the second, and a small management zone within a field in a third. No data were lost in transit in this case; what was lost was their intended meaning. This is more dangerous than a visible error message because the system may accept the file, complete its calculations, and produce apparently correct maps and reports whose underlying relationships are fundamentally wrong.
The same goes for familiar words like “yield,” “production,” “moisture,” and “farming date.” By “production” one system may mean the total weight before sorting, and another may mean the marketable quantity after excluding the wastage. “Moisture” may refer to the moisture of the soil, air or grain after harvest. The “planting date” may be the day the seeds were sown in the nursery, the day the seedlings were transferred to the field, or the beginning of the administratively recorded season. Interoperability is therefore not achieved merely because systems agree on file extensions or column names. Technical compatibility answers, ‘Can the system read the data?’ Semantic interoperability answers the harder question, ‘Does the system understand the data in the sense intended by their source?’ Four levels of compatibility Compatibility between systems can be conceptualized at four interrelated levels:
- Technical compatibility: The ability of two systems to communicate and exchange a file or message, such as sending a CSV file or responding via a programming interface.
- Structural compatibility: Agreement on the position, arrangement, and types of fields—that this column is a date, another a number, and a third the name of a crop.
- Semantic interoperability: Agreement on the meaning of fields, values, and relationships—what kind of date is meant, what the number measures, and which crop or field the name denotes.
- Operational and legal compatibility: Their agreement on what may be done with the data after it has been transferred, and who can correct it, publish it, or use it in a model or decision.
An exchange may succeed at the first level and fail at the other three levels. Therefore, it is not permissible to declare the link successful simply because the file was opened or the communication interface returned a successful response. The following table shows how simple fields can have different meanings:
| Field name | Possible meaning in a system | Possible meaning in another system | The consequence of confusion |
|---|---|---|---|
| Region | Governorate or administrative region | A zone within a field | Linking a local measurement to an entire administrative area |
| Planting date | Date of sowing | Date on which seedlings were transplanted to the field | Error in calculating plant age and growth stage |
| Production | Total weight at harvest | Marketable product after sorting | An unfair comparison between fields or seasons |
| Area | Area actually cultivated | Total holding area | Error in calculating yield per unit area |
| Soil moisture | Reading at a depth of 20 cm | Average readings at multiple depths | Recommendation based on unequal measurements |
| Treatment | Every spraying operation | Chemical treatment only | No record of biological or mechanical control |
| Price | Farm-gate price before transport | Wholesale price after packing and transport | A misleading financial conclusion |
| Status | Crop condition | Record review status | Interpreting ‘approved’ as a description of the plant rather than of the data |
Data dictionary: written agreement on meaning Every system needs a data dictionary that does not limit itself to the name of the column, but also explains its meaning and limitations. A good field description includes:
- A clear name that the user understands.
- A fixed identifier used by systems, which does not change when the name is translated.
- A definition that explains what a field represents and what it does not.
- Type of expected value: number, date, text, location, or selection from a list.
- The unit and method of measurement, if the field is a measured quantity.
- Allowed values and the meaning of each value.
- The level of place and time to which the information belongs.
- Whether the field is original or derived from other fields.
- Rules for dealing with absence and non-applicable values.
- The entity responsible for defining the definition and approving its change.
- The identification version number and its effective date.
If the field is “yield,” the dictionary should clarify whether what is meant is wet or dry weight, whether the measurement is before or after sorting, what is the unit of area, and whether the cultivated or harvested area is calculated. Without this, the system might compare two numbers with the same name but not measuring the same result. A fixed identifier is more important than a variable name Names change depending on language, region and spelling, but the reference identifier should remain constant. The crop may appear in its Arabic name in one file, in its Turkish or English name in another file, and in a local abbreviation in a third record. If the association is based on word similarity alone, the system may combine two different crops or separate two nouns that refer to the same thing. The solution is not to delete local names, but to associate them with a referring entity with a fixed identifier. The entity maintains the official name, scientific name, local synonyms, and areas of use, along with a degree of confidence in each match. When a match is ambiguous, it remains in a review state rather than being automatically approved by the system. This applies to fields, farms, devices, laboratories, varieties, growth stages, diseases, pests and practices. A stable identifier allows changing the apparent name or correcting the translation without breaking historical ties or producing new versions of the same entity. Relationships are part of meaning It is not enough to know that a value concerns ‘soil moisture’. We must also know its relationship to the field, zone, depth, device, time, and crop. Agricultural data do not live in separate lists, but in a network of relationships: This reading was taken from this sensor, in this sector, inside this field, at this depth, during this season, and it was related to this crop and at this stage of growth. If any of these relationships is lost, the number may remain while its validity disappears. Combining measurements from two depths may produce an average that represents neither soil layer; linking a crop to a similarly named but different cultivar may produce a recommendation unsuitable for the cultivar actually grown. Versioning silently prevents the past from being changed Dictionaries and schemas change over time. One concept may be divided into two more precise concepts, the definition of “marketable production” may be changed, a new unit may be added, or a relationship between a variety and a crop may be corrected. These changes are normal, but they become dangerous if they occur without a clear release. The system should know which version of the dictionary or schema was used when each batch was imported. If the definition later changes, old records are not silently reinterpreted. Instead, a documented migration process records what changed, which records were affected, and whether earlier results must be recalculated. This can appear to the auditor in this form: Version used when importing: Crop Dictionary 2.1 Current version: Crop Dictionary 2.3 Influential change: Separating a local name that was linked to one crop into two possibilities that need to be reviewed Action: Affected records were blocked and not automatically modified The error at the boundaries may equal the model error The Digitization of Plant Production review describes an interconnected chain of sensing, data, context, decision and action [SRC035]. In this series, the risk does not occur within the model alone; it may start at the boundaries between a device and a server, or between a file and a platform, or between two dictionaries. A timezone may be lost in transfer, causing a measurement to appear at the wrong hour. A cultivar name may be shortened on export and linked to another cultivar on import. A blank cell may become zero, measurements from different depths may be combined, or the unit definition may be lost while the number remains. The final result may look precise, but mathematical precision cannot restore meaning lost along the way. Therefore, linking must be tested with real examples, not just with field names. One system exports a small set of normal, ambiguous, and missing instances, then the other system imports them, and the team compares what arrived with what was intended. The test must include values, units, dates, synonyms, relationships, quality conditions, and rights, not just the number of rows. The role of artificial intelligence in matching AI can suggest that a column named “Moisture” likely refers to soil moisture, that a local name matches a known crop, or that two units are convertible. But it should not turn the possibility into a silent fact. Every proposal needs a degree of confidence and an explanation of the indicators on which it is based, and ambiguous cases go to human review. Good semantic interoperability does not force systems to use the same words, but rather makes each system able to understand what the other meant, see where there is disagreement, and refrain from linking when evidence is insufficient. When this is achieved, the transfer of data becomes a transfer of meaning, not a transfer of numbers alone. But understanding the meaning does not guarantee that the data is valid for every decision. A record may be clearly defined, united, and sourced, but then it may be outdated, incomplete in coverage, or biased toward particular locations. Here we move from the question “What does the data mean?” To the question: “Is its quality sufficient for this purpose?”
7. Quality Fit for Purpose
A low-resolution mobile-phone image may be sufficient to alert a farmer to inspect a patch in the field, but insufficient to determine the dose for a site-specific treatment. One sensor reading may be adequate for tracking a moisture trend in a small zone, but it does not represent an entire heterogeneous field. A weekly average price may suit a general report but not a selling decision in a market that changes within hours. This is why there is no absolute “high quality” data. There is data that is appropriate or inappropriate for a specific question, at a known place, time, and error consequence. Quality is not a characteristic that sticks to the file once, but rather a judgment that changes with change in use. The same record may be valid for viewing in a dashboard, valid for training a model, valid after human review for reporting, and prohibited from entering automated control. Therefore, the platform should not just say that the data is “good” or “bad,” but rather explain: good for what purpose? Within what limits? What is not permissible to do with it? Quality dimensions Data quality is examined across multiple dimensions, not a single number that hides strengths and weaknesses:
| Quality dimension | The question it answers | A simplified agricultural example |
|---|---|---|
| Completeness | Did all required fields and records arrive? | Moisture readings are present for most hours of the day, with missing hours identified |
| Coverage | Do the data represent the necessary space, time, and categories? | Distribute the sensors to different areas of the soil, not just collect them near the electricity source |
| Accuracy | How close is the measurement to the true value? | Comparing a moisture sensor with a reference measurement or reliable field test |
| Bias | Do the errors tend in a certain direction? | A device that gives consistently higher readings in salty soil |
| Currency | Does the value still describe the present condition? | A moisture reading from ten minutes ago may be suitable for monitoring, while one from a day ago may be unsuitable for immediate irrigation |
| Frequency | Is the measurement repeated often enough to capture change? | Measuring cold-room temperature every five minutes rather than once a day |
| Consistency | Do the values and definitions agree between the sources? | The unit area or definition of yield does not differ between two reports for the same season |
| Traceability | Can the value be traced to its source and transformations? | Knowing the file, row, device, and transformation rule that produced it |
| Representativeness | Do the data include the different locations and groups that matter? | Training data are not confined to large, well-connected farms |
| Rights integrity | Do the licence and consent permit this use? | Data suitable for internal analysis but not approved for training a commercial model |
| Resilience to interruption | Can data collection and operation continue when the network is weak? | Store readings locally, then synchronise them without losing their order |
| Recoverability | Can data be recovered after a failure or accidental deletion? | A backup tested by an actual restoration, not merely a promise that a copy exists |
One dimension cannot compensate for the absence of the rest of the dimensions. The data may be very accurate for three farms, but it is not representative of the region. It may be complete and up-to-date, but its source is unknown. It may be well documented, but the contract does not allow it to be used in training. In each case, the result is different: data that is useful for one purpose, and misleading or unlawful for another. Card quality, no stamp It is useful for every record or dataset to carry a brief ‘quality card’ that states its status, rather than a generic stamp saying ‘approved’. The card may display: Proposed purpose: Monitoring the daily moisture trend Coverage: Three of the four sensors are working Currency: Latest reading taken 12 minutes ago Calibration: Two devices within the period, and the third needs to be checked Representation: The western sector is not covered Decision: Valid for warning and review, and not valid for automatic irrigation operation This card is more informative than a single percentage such as ‘82% quality’. A composite score may conceal that data are current and complete but do not cover the area where the decision will be applied. It may also permit an unacceptable trade-off: a reading's currency cannot compensate for lack of permission to use it, nor can file completeness compensate for failed calibration in a high-risk decision. Different thresholds for different uses The higher the error penalty, the higher the quality requirements. The uses can be divided, in practical terms, into levels:
- Display and exploration: Incomplete or low-confidence records may be displayed, provided that warnings are clearly visible.
- Analysis and reporting: Needs consistent definitions, known coverage, and an auditable conversion history.
- Training models: In addition, it requires appropriate representation, bias checking, separation of training and testing data, and valid usage rights.
- Operational recommendation: Requires greater currency, local context, stated limits, and risk-based human review.
- Automated control: Requires the highest degree of verification, independent sources where possible, safety limits, and immediate ability to shut down and return to safe operation.
A low-confidence reading may be useful to a reviewer because it draws attention to an area requiring inspection, while remaining excluded from a published report, a training dataset, or an operating command. This does not waste the data; it uses them only at the level their quality permits. Completeness does not mean representation The file may be full of fields and not represent the reality in which it will be used. If plant disease images are compiled from clear leaves taken by professionals in good light, the images may be technically excellent, but they do not represent the shaky phone images, dusty leaves, or harsh lighting the user encounters in the field. Yield data may be complete for large farms using digital systems yet missing for smallholdings, paper-based records, and poorly connected areas. A model then learns from the group easiest to record, not from the entire agricultural community. Quality therefore asks, ‘Who appears in the data?’ and ‘Who is missing?’, as well as how many fields are complete. Quality changes with time A record that passes checks when collected may lose its validity later. A device's calibration may expire, its location may change, the item being observed may change, the laboratory method may change, or the platform may adopt a new data-dictionary version. A model may remain accurate for one season and then deteriorate because of unusual weather or a change in agricultural practice. This is why quality needs to be monitored, not a single review. Absence rates, device skew, variability in distributions, frequency of corrections, and differences between sites and categories are monitored. When an indicator exceeds a known threshold, the quality status is lowered, the data is quarantined, or a new validation is requested. Failure in quality should change behaviour Quality knowledge that does not result in action is worthless. If the data is low confidence, the system must know what to stop and what remains allowed. Actions can be:
- Display a clear warning to the user.
- Submit the log to a checklist.
- Request an alternative measurement or additional image.
- Preventing publication or training.
- Exclude the record from the automated decision.
- Return to a more conservative operating base.
- Withdrawal of results that were based on a batch proven to be corrupt.
These rules must be known before a problem occurs, not invented after damage appears. Quality is not a report describing data from a distance; it is part of the operating logic. Professional quality question Instead of asking “Is this data good?”, we ask: What decision will you support? What is the consequence of the error? What is the minimum level of completeness, accuracy, timeliness and representativeness of this decision? What does system do when this limit is not met? Can the user see the reason and object to it? When quality is tied to purpose, it's possible to use lean data for low-risk alerting without giving it undue authority, and to protect sensitive decisions from seemingly tidy but inadequate data. However, applying these rules requires the ability to go back from each result to its origin, and know what has changed along the way. This is the role of provenance and transformation log.
8. Provenance and the Transformation Log
An agricultural system may display a forecast that a field will yield 6.4 tonnes per hectare. The number is clear, but the professional question does not begin there; it begins with what lies behind it. Which field and season? What data entered the calculation? Were historical yields measured or estimated? Which model version produced the forecast? Were area units changed, or were records excluded during processing? If the system cannot answer these questions, the number becomes an orphaned result. It can be displayed and used in a decision, but cannot be verified, distinguished from an earlier version, or withdrawn if a defect is found in one of its sources. Provenance is the ability to trace information back to its origin, and then trace everything that happened to it from the moment of collection until it appears in a report, model, or recommendation. The transformation log is the part that explains what changed, by what rule, at what time, by which person or service, and why. From the result to the field The auditor should be able to move in the reverse direction along a clear chain: Recommendation or forecast → model version → features used → normalised records → raw values → file, device, or laboratory → field or sample from which the observation began. This does not mean that the farmer sees this entire technical chain on every screen. The interface can display an understandable summary, while the details remain available to the technician, auditor, or auditor when needed. The table shows what each stage should answer:
| Stage | The question it must answer | Simplified example |
|---|---|---|
| Result | What was issued, to whom, and when? | Yield forecast created on August 15 for a specific farm |
| Model or rule | Which version produced the result? | Yield Model, Version 3.2 |
| Derived inputs | What variables were used? | Total rain, average temperature, and vegetation growth index |
| Normalised records | How did the values become computable? | Converting area from feddans to hectares and standardising the date |
| Raw values | What actually arrived before transformation? | The text ‘5 feddans’ and the rain reading as transmitted by the device |
| Direct source | Where was the value found? | ‘2026 Season’ file, ‘Field 7’ sheet, row 148 |
| Field source | Who or what created the observation? | Harvester scale, weather station, or laboratory |
| Agricultural context | Which reality do you belong to? | Tomato field, eastern sector, 2026 season |
What is saved when you import a file? When importing a tabular file, it is not enough to save the final values in the database. Depending on need and risk, retention should be made to allow reconstruction of what happened:
- File ID and name at the time of receipt.
- A digital fingerprint detects changes in its content.
- The source of the file, the date it was uploaded, and who uploaded it.
- Sheet name, row and column number.
- Column header as originally stated.
- raw value before cleaning or converting.
- The language and local setting used in its interpretation.
- The original unit and the normalised unit.
- The rule used in the conversion and its version number.
- The entity to which the record is linked and the degree of confidence in the link.
- Quality flags, warnings and errors.
- Review and approval status.
If a column is headed ‘Production’, the system does not infer its meaning from the name alone. It retains the original header, file, and sheet, then records that a reviewer or classifier linked it to the concept ‘marketable weight after sorting’ at a stated confidence level. If it later emerges that total weight before sorting was intended, every affected record and result can be identified. Each conversion is an independent event Cleanup should not be a command string that disappears after execution. Each significant conversion is recorded as an event that contains:
- The input that was used.
- The generated output.
- Operation type: unit conversion, date interpretation, entity linking, value substitution, or record exclusion.
- The rule, code, or model and its version.
- The person or service that carried out the operation.
- Implementation time.
- The reason for the operation.
- Degree of confidence or warning.
- Who reviewed or approved the result, if it required review.
Thus, the system can say, for example: The value reached 0,24 from the Turkish setup file. The comma was interpreted as a decimal point, the value was converted to 0.24, and then displayed to the user as 24%. The conversion was carried out in Local Number Base, version 2.1, and approved by the Data Officer in Batch No. 46\. This statement is more valuable than the bare figure 24%, because it makes clear that the result was not the original text and allows recalculation if the interpretation rule changes. Provenance Is a Network, Not a Footnote A single reading may enter more than one product: an irrigation report, a dashboard, a forecast model, and an alert to the farmer. A single report may be based on thousands of readings from multiple devices and files. So provenance is more like a grid of attributions than a note at the bottom of the page. The system should be able to answer in both directions:
- Back: Where did this result come from?
- Forward: What reports, forms, and decisions used this record?
The second question is crucial when discovering the error. If a laboratory is found to have used an incorrect calibration method in a batch of soil analyses, correction should not be limited to adjusting the laboratory schedule. You should inventory the fields, recommendations, reports, and forms that relied on that payment, and then determine what should be recalculated, withdrawn, or notified to its users. Why are the two versions different? A farmer may ask why the yield forecast changed from 6.4 to 5.9 tonnes per hectare although the field itself did not change in one day. A system with sound provenance can answer: rainfall data arrived late, the field area was corrected, the model version changed, or a sensor shown to have drifted was excluded. As for the system that does not keep a transformation log, all it has to do is say that “the system recalculated.” This is an answer that does not allow for learning, objection, or assigning responsibility. The explanation for the difference can look like this: Previous result: 6.4 tonnes/ha Current result: 5.9 tonnes/ha Reason for the change: Corrected rain data for three days arrived Unchanged: Field area, variety, and yield calculation method Model version: Unchanged Review status: The update is automatic, and appears to the specialist before the report is approved Provenance Is Not Unlimited Retention Saving a trace does not mean copying everything forever or making details available to every user. Raw logs may contain sensitive personal, business, or location data. So provenance itself is subject to access permissions, retention periods, and deletion policies. It is possible to preserve evidence that a deletion occurred, its reason and date, without retaining the content that should have been deleted. Some details can be hidden from the general user while keeping them for a party authorized to audit. The goal is to combine accountability with respect for access limits, not use tracking as an excuse for unlimited collection. The value of provenance in daily work These details are not an archival burden separate from agricultural work. It allows:
- Reconstruct a result that users disagreed with.
- Determine why a report changed between two versions.
- Withdraw wrong payment without deleting proper data.
- Find out which devices or files need to be rescanned.
- Evidence that the farmer supplied or corrected particular data.
- Distinguish true measurement from estimated value.
- Review whether the data was used within the permitted purpose.
- Rerun the analysis after correcting a rule or model.
A traceable system does not promise that errors will never occur; it ensures that an error does not become a mystery of origin that cannot be reversed. When the source and history of every number are known, the associated rights can be discussed: who may see, correct, or transfer it, and who may leave the system without losing their agricultural history.
9. Data Rights, Portability, and Exit
A farmer may use a platform for years, recording field boundaries, soil-analysis results, irrigation and fertilisation schedules, crop photographs, production costs, machinery history, and seasonal notes. When the farmer decides to move to another service, they may discover that the only downloadable output is a PDF report: readable, but impossible to import into an alternative system and unable to reconstruct the relationships among fields, seasons, and readings. At this moment it turns out that the presence of an “Export” button does not mean portability, and that the ability to see data does not mean the ability to recover it or use it outside the platform. The data may be tied to the farmer's account in theory, but it is locked into a format, identifiers, or links that only work within the vendor's system. The OECD analysis highlights ownership, access, trust, and interoperability from farmers' perspective [SRC032]. Yet the question ‘Who owns the data?’ is not sufficient on its own, because ownership comprises several rights and powers that may be distributed among the farmer, worker, laboratory, platform provider, and public body. The more accurate question is: who gets to do what, with what data, for what purpose, for how long, and under what oversight? Rights are a bundle, not a single word Data rights can be broken down into practical questions:
| Right or power | Practical question | What should the system provide? |
|---|---|---|
| Access | Who can view the data? | Clear role-based permissions and a log showing who accessed the data and when |
| Collection | Who may create or import the record? | The identity of the observation's creator or the file's uploader, together with the basis for collection and any required consent |
| Use | Which purposes are permitted? | Link each dataset to defined purposes and prevent uses outside them without fresh permission |
| Correction | Who can object or correct an error? | A correction path that preserves the original, the amendment, the reason, and the identity of who approved it |
| Sharing | May the data be sent to a third party? | Recipients, purpose, duration, and whether transfer is continuous or one-off |
| Training | May the data be used to train a model? | Separate, explicit consent rather than consent hidden inside general permission to operate the service |
| Publication | May the data or results be displayed publicly? | Rules for anonymisation or aggregation, and review of what might reveal the farm or business |
| Portability | Can users download their data in a usable format? | Export raw and normalised values, relationships, and the transformation log in documented formats |
| Delete | What can be deleted, when, and what are the exceptions? | Verifiable implementation, indicating what was deleted, what remains, and why |
| Objection and suspension | Can the user stop new usage? | A means to withdraw permission and prevent future use of derivative processing in accordance with the policy |
| Exit | What happens when the contract ends? | Transition plan, download period, revocation of access keys, and deletion of unauthorised copies |
This decomposition prevents misleading answers. A vendor may say, ‘The data belong to the farmer’, while reserving a contractual right to use them for training commercial models, giving the farmer a file that cannot be imported, or deleting the account before the history has been transferred. Declared ownership is insufficient without a practical ability to access, correct, transfer, and refuse uses of the data. Consent to the Service is not consent to all use A system may need moisture data to make an irrigation recommendation, but that does not automatically mean it is allowed to be sold, combined with commercial data, used to train a generic model, or publish maps that reveal farm activity. Each new use needs a clear purpose, permissible basis, and information that the user understands. Consent should be specific and returnable, not a page long request to be accepted all at once. The farmer must know:
- What data will be collected?
- Why does the system need it?
- Which work or service stops if the farmer does not provide them.
- Who will receive it inside and outside the organization?
- Will it be used to train a model or develop another product?
- How long will it last, and where will it be stored.
- How can the farmer correct or download them, or request their deletion?
- What happens if the farmer withdraws consent or terminates the contract?
Denial of secondary employment should not turn into deprivation of a primary function for which it is not needed. Free consent sets realistic alternatives, and does not make universal acceptance a price for using a necessary service. Portability Is More Than Downloading a File For data to be portable, the export must not be limited to a readable report. Depending on the type of data, the user needs:
- raw values as arrived.
- Standardized values with units and conversion rules.
- A dictionary explaining field names and their meanings.
- Identifiers and relationships between farms, fields, seasons, devices and observations.
- Locations and boundaries in a spatial format that can be used when needed.
- Quality, confidence, and approval states.
- Basic history of corrections and conversions.
- Source, file, sheet and row information when importing.
- Licenses and restrictions associated with reuse.
- Clear documentation of the format and import method.
CSV format may be fine for a simple table, but it alone is not sufficient to preserve a network of relationships, geographic boundaries, or history of transfers. The export may require more than one file and format, with an index explaining how the parts are related. The important thing is that an independent system can reconstruct meaning, not just open files. Test out before you need it It is not enough for a contract to state that ‘the user has the right to export their data’. Exit capability is an operational capacity that must be tested periodically, like a backup. A simple test can be run annually or after a major update:
- Export a sample representing fields, seasons, readings, images, and relationships.
- Verify the presence of raw, standardized values and metadata.
- Import the sample into a standalone tool or demo environment.
- Compare number of records, values, units and dates.
- Ensure that the field links to the season, device, and source remain.
- Check whether local names, quality statuses and rights have been transferred.
- Document what was lost or changed, and specify a responsible person and a date for repair.
This testing early detects problems that do not appear in contracts: meaningless off-platform identifiers, undocumented modules, unlinked image files, or key fields not covered by the export. What happens when the contract is terminated? Checkout is not limited to downloading data. The plan must answer the operational, security and rights questions:
- How long does it take to download before the account is closed?
- Does the system continue to function securely during the transition?
- Who helps interpret or import the formula?
- What happens to the hardware communication switches and programming interfaces?
- Are field devices still usable locally?
- When are the vendor's and its employees' permissions revoked?
- What copies are deleted, and what copies must be legally or contractually retained?
- What happens to the models trained on the data, or to the results derived from it?
- How does the vendor prove that the deletion or move is complete?
It may not be possible to erase data from a model that was trained on it in the same way as deleting a row from a database. This issue should therefore be discussed before training, and the withdrawal policy, model versions, what can actually be implemented and what cannot be guaranteed should be determined from the beginning. Exit is part of agricultural sovereignty When farmers can move their records to another system, the vendor becomes a replaceable service provider rather than the sole gateway to the farm's history. Cooperatives and institutions gain bargaining power, the cost of changing systems falls, and innovation becomes possible without rebuilding the record from scratch. However, if limits, readings, seasons and corrections only work within one platform, the accumulation of data from an asset serving the farm may become a restriction that prevents it from leaving. This is why portability and exit are designed at the beginning of the system and nodes, not at the end. However, the right to access and transfer is not complete without protecting data and functions from tampering and disruption. An exportable but poorly secure platform may give an attacker the same access a legitimate user needs. Hence cybersecurity becomes part of agricultural safety, not an isolated technical function.
10. Cybersecurity as Agricultural Safety
In a traditional desktop system, a compromised account might lead to a file leak or service disruption. In a connected farm, the effect may be transmitted from the screen to plants, animals and food. Changing the watering duration, closing an air vent, disabling store cooling, or adjusting a feed dose can cause physical damage before the user discovers that the problem started with a digital attack or compromised setup. This is why agricultural security is not measured by the number of strong passwords alone. A review of cybersecurity in smart agriculture emphasizes that the expansion of connectivity, devices, and services opens up a broader risk surface [SRC030]. Risks include data theft, alteration, disruption, device spoofing, hijacking of operating functions, and provenance logs. The basic security question is not only “Can an unauthorized person get in?”, but also: What can they do if they get in? What is the outcome in the field? How long does it take to detect change? Can the operator stop it and return to a safe position? Three properties that must be protected Security protects three interconnected dimensions:
- Confidentiality: Only those with authority can see the data. This includes field locations, prices, costs, and worker and customer data.
- Integrity: Data and commands must not be changed without detection. A manipulated moisture reading or forged irrigation command may be more dangerous than a leaked report.
- Availability: The basic function remains available when needed. Stopping the ventilation or cooling system at a sensitive time may be more serious than temporarily losing access to the reporting panel.
It is not enough to protect one of these dimensions and neglect the rest. Data may be well encrypted but unavailable during a crisis, a service may be available but an attacker can modify commands, or commands may be intact but the platform exposes sensitive business information. Classification of assets according to the consequence of the defect Not all data and functions need the same level of protection. It is best to classify them according to what would happen if they were detected, changed, or stopped:
| Asset type | Example | Possible consequence | Required protection |
|---|---|---|---|
| General information | Published instruction manual | Limited damage if copied | Protection from unauthorized change |
| Normal operational data | Machine maintenance record | Late or incorrect maintenance decisions | Permissions, backups, and change history |
| Personal or business data | Workers' wages, production costs, and selling prices | Harm to privacy, negotiation or reputation | Restricted access, encryption, and a clear retention policy |
| Provenance and quality data | Transformation log, calibration, and review | Difficulty establishing error or withdrawing results | Tamper-evident log and independent review |
| Sensitive operational control | Pump operation, ventilation or cooling | Direct damage to crops, animals or food | Network separation, security boundaries, strong authentication, and local control |
| Safety critical function | Emergency stop or dangerous temperature alarm | Significant damage if it is disrupted or delayed | Independent path, periodic testing, and manual operation capability |
This classification prevents effort from being spent equally on everything. Protecting a public page is not the same as opening a valve, and accessing a historical report is not the same as accessing a live cooling system. Where does the risk enter? The attack surface in connected agriculture is not limited to the central server. The danger may enter through:
- Sensor with default password.
- Old communication portal that has not received updates.
- A lost phone still has an active login session.
- Shared computer on the farm.
- A tab file or external volume is infected.
- A supplier account has broad support authority.
- An exposed API or insecurely stored key.
- Unreliable software update.
- A wireless network that combines visitor devices with control devices.
- A fraudulent message asking the worker to enter the password.
- External instructions or content that attempt to mislead an AI tool connected to the system.
Each ring needs a clear administrator, update cycle, and detection method. It is not correct to assume that the device is safe because it is small or located in the field, or that the vendor will address the risk without agreement and testing. Least Privilege and Separation of Duties Each user, device or service gets the minimum permission necessary for its operation. An operator recording an observation does not need to delete an entire season, an analyst reading historical data does not need to operate a pump, and a reporting service does not need to hold a ventilation control switch. Control networks and functions should also be separated from less trusted interfaces. If your marketing board or email account is attacked, the attacker will not find a direct route to the irrigation devices. Sensitive changes, such as adjusting automation limits or disabling an alarm, will preferably require additional verification or approval from a second person depending on the level of risk. Core practices include:
- A separate account for each user, instead of a shared password.
- Multi-factor authentication for sensitive accounts.
- Revoking the authority immediately after changing the role or expiring the contract.
- Rotate access keys and secrets and do not place them in exposed files.
- Updating devices and services according to a plan and testing.
- Record login attempts and critical changes.
- Review vendor permissions and remote support.
- Sign updates and verify their source when possible.
Safe failure is more important than continuing at all costs Farming doesn't stop when the Internet goes out. So the system must know what to do when it loses connection, receives an unreliable command, or two critical readings differ. Safe mode doesn't always mean stopping everything. Stopping ventilation inside a greenhouse may be dangerous, as can continuing watering without limits. The safe mode for each job is determined based on plant, animal, environment and season, and may include:
- Continuation of local control within governorate borders.
- Reject new remote commands while keeping the last trusted setting for a specified period.
- Determine the maximum time or amount of operation.
- Require human verification before exceeding a sensitive limit.
- Local alarm triggering that does not depend on the cloud.
- Provide clear and secure manual switch.
- Record what happened to review after the connection is restored.
The operator must be trained to stop the automation without losing basic function. A kill button that the worker doesn't know where it is, or a manual procedure that hasn't been tested in years, is not a true emergency plan. Backup does not equal restore The platform may claim that it creates daily backups, but the real value appears when you try to restore. The copy may be incomplete, encrypted with the same key that was lost, save tables without files and relationships, or take more time to recover than is feasible. Therefore recovery is tested periodically on a separate sample or environment. It is measured:
- What data and functions can be restored.
- How long did the return take?
- How much time was lost between the last version and the crash?
- Do the relationships, proofnance and permissions remain correct?
- Can the team perform the action when the vendor is absent.
Incident response begins before it occurs When an unusual watering issue or sensitive setting change is detected, the team needs a plan, not improvisation. The plan specifies:
- Who has the right to stop the function or isolate the device.
- How to preserve evidence and records without continuing damage.
- What is the manual or local alternative to continue the necessary work?
- Who should be informed on and off the farm.
- How the data is examined and the decisions that may have been affected.
- When reconnection and operation are allowed.
- How are the incidents documented and controls updated afterward?
An AI system should not send a critical command to a machine just because the wording of the recommendation is confident. Commands go through verification rules, trigger limits, and authentication, and high-consequence decisions are subject to independent human or automated review. The soundness of the language does not prove the soundness of the action. Security is part of agricultural confidence A safe system is not a system that promises the impossibility of hacking, but one that reduces its chances, limits its impact, detects it, maintains secure functionality, and can explain what happened. When data is protected from tampering, control functions are separated, and manual recovery and shutdown are tested, security becomes an extension of irrigation, cooling and feeding safety, not a technical addendum after the project is complete. However, protecting data is not limited to preventing an attacker from accessing it. A system might be technically safe, but then compromise local knowledge in an apparently legitimate way: extract it, standardize it, attribute it to its dictionary, and erase its owners and context. Here comes another risk that passwords don't address: excessive standardization and cognitive acquisition.
11. Local Knowledge and Excessive Standardisation
A farmer may describe a change in a plant with a word that does not appear in any official dictionary, or relate the direction of a local wind to the time of spread of a pest, or distinguish between two soil conditions that a system classifies under one classification. This knowledge may not come in the form of a number or a lab report, but it is the product of long seasons of observation, work, and collective memory. Systems need standards to link, search, and compare data. Yet a standard may change from a bridge between meanings into a tool that erases them if everything outside the central dictionary is treated as error or noise. Good standardisation makes difference intelligible; excessive standardisation makes it invisible. The local name is not necessarily misspelled A local name may have one of several roles:
- Synonymous with a well-known concept.
- A broader name that brings together several cases that the specialist differentiates between.
- A narrower name that describes a precise local condition.
- A description of an apparent symptom, not a diagnosis of its cause.
- A name whose meaning varies from one village or region to another.
- Knowledge not yet included in the official dictionary.
So system should not choose the closest standard word and then delete the original. It retains the local text, language, region, context and owner of the information, then links it to the reference concept with a clear relationship: “synonymous”, “broader than”, “narrower than”, “related to”, or “possible match that needs to be reviewed”. The table shows the difference between meaning-preserving monotheism and excessive monotheism:
| Case | Excessive standardisation | Meaning-preserving treatment |
|---|---|---|
| Local name for a growth stage | Replaced directly with the standard stage name | Preserve the local name and link it to the probable stage, recording region and context |
| A farmer's description of plant symptoms | Converted into a confirmed disease diagnosis | Preserve the description as an observation and leave diagnosis to an independent verification pathway |
| Local indicator of weather change | Treated as unstructured text and deleted | Record it with time, place, attribution, and then study its relationship to official measurements |
| Local soil classification | Collapsed into one broad category | Document the characteristics farmers intend and compare them with the scientific classification |
| Cultivar name with several spellings | Automatically create a new entity for every spelling, or merge them without review | Link spellings to a reference identifier after verification, preserving both registered and local names |
| Inherited practice | Summarised under a general word such as ‘traditional’ | Describe its steps, timing, conditions, and who has the right to share it |
The reference entity is a bridge, there is no alternative A reference entity with a stable identifier can be created for each crop, cultivar, disease, or practice and linked to its local names. The reference entity should not swallow the original wording. Its purpose is to support search and linking across languages and systems while preserving every designation in context. The term local may be associated with two possible concepts depending on the region, or its meaning may change between two generations of farmers. A mature model allows for multiple parallel maps, and records who proposed each map, who reviewed it, and the degree of confidence in it. The system does not force the true difference to a single answer for the sake of database convenience. The log can appear like this: The text as the farmer said it: Preserved in audio and writing Language and region: specified The meaning explained by its author: Description of a condition that appears after a certain pattern of wind Closest reference concept: Plant stress symptoms — possible match What the record does not prove: It does not prove a specific disease or cause Use condition: Available for local research, not approved for treatment recommendation The documentation must be kept by its owner Local knowledge does not arrive in the system out of thin air. They are provided by individuals and communities, and may be linked to identity, occupation, reputation, and economic resource. So system should log, depending on consent and context:
- Who provided the knowledge or which community you belong to.
- How does the owner want it to be attributed?
- What purposes are allowed?
- What parts may not be published or located.
- Is it permissible to translate it, summarize it, or train a model on it?
- How long should it be kept and how can permission be withdrawn?
- What benefit will accrue to those with knowledge?
Some knowledge may be collective, with no single person authorised to permit its use on everyone's behalf. It may also be sensitive because it reveals the location of a rare resource or an economically valuable practice. In such cases, technical consent inside an application is insufficient; governance requires appropriate community representation and clear access boundaries. Benefit is not a thank you at the margin The extraction process does not become fair just because the name of a community is mentioned in a report. If the knowledge is used to improve a paid product, model or service, the form of the benefit should be clearly discussed. The benefit may be:
- A financial return or a share in the revenue according to the agreement.
- Guidance service or tools that give back to the community.
- Local training, infrastructure and data management capacity.
- Free or preferential access to the results.
- Participation in decisions about updating and publishing.
- A clear lineage that protects those with knowledge from erasure.
There is no single model that fits all cases, but the rule is that the benefit is achieved before extraction and use, not after the knowledge becomes part of a product that is difficult to separate from. Artificial intelligence risks on local knowledge A language model can translate local accounts, extract terms, and suggest links to a scientific dictionary. At the same time, however, it may:
- Turns a potential description into a definitive fact.
- Combines two similar terms and erases the difference between them.
- The name is translated linguistically and loses its agricultural meaning.
- Paraphrases knowledge without attributing it to its owner.
- Displays sensitive information outside of the context in which it was permitted.
- The term that is more present in its data is preferred over the more precise local term.
Extraction, translation, and linking therefore pass through human review by people who know the language, agriculture, and local context. The system must preserve the original wording, the model's proposal, and the reviewer's decision, rather than allowing automated output to replace the testimony of the knowledge holder. Participation of knowledge holders in governance It is not enough for the project to consult farmers once at the beginning. The local dictionary, term maps and access policies need to be reviewed periodically. Participation may include:
- Verification sessions for names and definitions.
- The ability to correct or object to the link.
- Interfaces support local language, voice and offline working.
- A committee or representatives review new uses.
- A report explaining where the knowledge was used and what results emerged from it.
- A means to withdraw material or restrict its publication in accordance with the agreement.
Success is not measured by the number of terms that the system has “cleaned,” but rather by its ability to retrieve knowledge in its language and context, link it to others without erasing it, and keep its owners able to understand, correct, and control. Standardisation is necessary for systems to communicate with one another, and local knowledge is necessary for them to speak meaningfully about agricultural reality. Mature governance does not choose between the two; it makes the reference entity a bridge while preserving origin, context, attribution, and benefit. After all these principles, a practical question remains: when purchasing or reviewing a platform, how do we know that it actually applies them?
12. Agricultural Data Platform Checklist
The platform might display a beautiful screen, real-time graphics, and a model that predicts yield or recommends irrigation. But the quality of the display alone does not reveal whether the raw value is preserved, whether the export is usable, what happens when the network is down, or who owns the operation of a connected machine. So the following checklist is used before purchasing, connecting or refurbishing, and then repeated after major upgrades and accidents. Questions are not answered with general statements like “system is secure” or “supports export,” but rather with evidence that can be examined: a screen, a test file, a log, a recovery test, or a clear contractual clause. How to use the list? The team scores for each question:
- Yes, with evidence: The capability exists and has been tested.
- Partial: Available in some cases or needs to be adjusted.
- No: Not available.
- Unknown: No evidence has been provided; in practice, treat it as unverified until evidence is supplied.
- Not applicable: With an explanation of why it does not apply.
The answers should not be collected into one score that hides the risks. Some failures are critical: a platform might get good answers on twenty items and then remain unsuitable for automated control because it does not have a safe outage mode. Therefore, each deficiency is linked to a consequence and use, and provisions are identified that prohibit deployment, training, or operation until they are remedied. First: the origin and meaning
- Does the platform keep the raw value as it arrived, instead of replacing it with the cleaned or corrected value?
- Can the auditor see the raw value, the standardized value, and the difference between them?
- Do you record the measured characteristic, unit, location, time, and method of measurement?
- Do you keep the time a measurement occurs separate from the time it arrives or is processed?
- Do you distinguish between zero, missing, not applicable, below the detection limit, device failure, and withheld value?
- Do you retain the language, local setting and text of the original when interpreting numbers and dates?
- Do you request a revision when a number, date, or unit has more than one meaning?
- Do you link local names to reference identifiers without deleting the original names and contexts?
Acceptable evidence: Opening a real record and showing the origin, interpretation, unit, time, and state of absence, not just a marketing document. Second: conversion and provenance
- Does each transformation have a reason, time, port, and versioned rule?
- Can a report or recommendation be traced back to the records, files, and devices on which it was based?
- When importing a file, are the identities of the original file, sheet, row, column, and header preserved?
- Does the platform use a fingerprint that detects changes in the raw file or record?
- Can it identify the results and models affected by a record later found to be wrong?
- Can a transformation or analysis be rerun after a rule is corrected, without collecting the data again?
- Does the platform explain why a result differs from an earlier version?
Acceptable evidence: Pick a number from a report, then practically trace it back to its source, or perform a test patch and show the affected results. Third: Quality and Approval
- Do you evaluate quality by purpose, rather than by a generic stamp like “good data”?
- Does the platform show completeness, coverage, accuracy, currency, representativeness, and rights status?
- Are the limits of the data visible to the user—for example, that a reading represents one point rather than the whole field?
- Can low-confidence records be quarantined instead of deleted or published?
- Do quality rules prevent an unsuitable record from entering training, publication, or automated control?
- Does an imputed value carry a status showing that it is an estimate rather than a measurement?
- Is quality reviewed over time as the device, location, season, or dictionary changes?
- Does the system know what to do when a quality threshold is not met: issue a warning, request a new measurement, or halt a decision?
Admissible evidence: Show a low-confidence record, then demonstrate that it is visible to the reviewer and does not fall into a prohibited use. Fourth: interoperability
- Is there a dictionary that explains the meaning of each field, its unit, and its allowed values?
- Do entities such as fields, crops, and devices use stable identifiers that do not depend on the display name?
- Do data schemas and dictionaries carry version numbers and effective dates?
- Are the relationships among field, sector, season, crop, device, and observation documented?
- Has data exchange with another system been tested for meaning and relationships, not merely row counts?
- Are proposed semantic matches and confidence scores displayed, with ambiguous cases sent for review?
Acceptable evidence: Export a sample to a standalone system and then re-import it and compare units, dates, entities and relationships. Fifth: Rights and access
- Does the user know who is seeing their data, for what purpose, and for how long?
- Are permissions role-based, with a record showing who accessed, changed, or exported the data?
- Can the farmer or another authorised party correct a record while preserving the amendment history?
- Does the system separate consent to operate the service from consent to train a model or share data?
- Can a user withdraw permission for future use, and does the platform explain what that means for derived data?
- Does the retention policy specify when data are deleted and identify any exceptions?
- Does the platform protect local knowledge from publication or training beyond the permission granted?
- Does it explain benefit and attribution when it uses knowledge supplied by individuals or a community?
Acceptable evidence: Review the permissions screen, access log, and consent texts, and execute a correction request or withdraw test permission. Sixth: Transportation and exit
- Can the user download raw values, consolidated values and historical record?
- Does the package include the data dictionary, units, quality states, and essential provenance?
- Do relationships, field boundaries, images, and files move with it, or only an isolated table?
- Do exports use documented formats that another system can read, rather than only a PDF report?
- Has an actual import test been performed in an independent system or tool?
- Does the contract specify how long downloads remain available and what assistance is provided at termination?
- Do essential devices and functions remain operational during the transition?
- Are the vendor's keys and permissions revoked on exit, and is evidence of the required deletion provided?
Acceptable evidence: A mini-exit experiment that includes exporting, importing, and checking that meaning and relationships remain. Seventh: Security and continuity of operation
- Are operational control functions separated from reporting, marketing and public access interfaces?
- Does every user, service, and device receive only the minimum privileges required?
- Do sensitive accounts use strong authentication, with access revoked immediately when a role changes or a contract ends?
- Are vendor and remote-support permissions reviewed, and are their sessions logged?
- Are sensitive changes recorded in a tamper-resistant log?
- Are there backups of the necessary data, relationships, files, and settings?
- Has restoration actually been tested, including recovery time and the amount of data that might be lost?
- Can the system continue to provide a safe level of service when the network or cloud service is interrupted?
- Are there limits preventing an unreasonable command from operating irrigation, ventilation, or cooling without constraint?
- Does the operator know how to stop automation and return to manual or local operation?
- Is there an incident plan specifying who isolates the device, preserves the logs, notifies those affected, and restores service?
Admissible evidence: Manual interruption, recovery or shutdown exercise in a safe environment, no theoretical description of the procedure. Eighth: Artificial intelligence and decision making
- Does the platform explain what the model produced and what came from a fixed measurement or rule?
- Do classification, linking, and imputation pass through an approved model whose version is recorded?
- Does the platform show confidence and its limits rather than presenting the answer as conclusive?
- Are low-confidence or high-consequence cases sent for human review?
- Is final publication or sensitive control based on unapproved outputs prohibited?
- Can the result be linked, to the extent permitted, to the training data or to the inputs and rules that produced it?
- Is model quality monitored after changes in season, region, variety, or data type?
- Can the user contest or override the recommendation and record the reason?
Acceptable evidence: A test case in which the model produces a low-confidence result, followed by evidence that the system quarantined it or referred it for review instead of acting on it. Gates that may not be crossed Not all questions are equally effective. There are conditions that should prevent use until treated:
- There is no raw copy to refer to.
- The unit, property, entity, or time in data that will enter a decision is not known.
- The system does not distinguish between zero and missing value.
- The result cannot be traced back to its source and underlying conversions.
- There is no clear authority allowing use, training or participation.
- Users cannot export their data in a usable format.
- Low-confidence records enter publication or control without review.
- There is no safe way to turn on or off when the connection is lost.
- A single low-trust account can access sensitive operational control.
- Recovery, egress, or incident response have not been tested.
Having one of these vulnerabilities does not always mean the platform will be rejected for every purpose, but it does specify what should be prevented. A platform may still be suitable for displaying general information, but not for training a model, running a machine, or keeping a long-term record. Who answers and when? The technology team cannot answer the checklist alone. Farmers or their representatives, agricultural specialists, data officers, security officers, operational users, and legal or contractual authorities participate as needed. The question about measurement depth is agricultural; the question about export format is technical; the question about permission for training is contractual; and the question about manual shutdown is operational and security-related. The review is repeated:
- Before purchasing the platform or signing the contract.
- Before connecting a new device or system.
- Before using the data to train a model.
- Before moving from presentation to recommendation or control.
- After a major update or change in vendor.
- After an accident or detection of a wrong data batch.
- Periodically within the governance and quality review.
The checklist is not an examination that awards a score; it is a tool for making assumptions visible. If evidence cannot be supplied, the field should not be filled with optimism. ‘Unknown’ is recorded, together with the responsible party, the risk, and a due date. Declared ignorance can be addressed; unsupported confidence passes silently into the decision.
Chapter Summary
Agricultural sovereignty begins not with a more complex model, but with a record capable of telling its own story. A number without a unit, place, time, or source may enter the fastest algorithms, yet remain poor in meaning. An observation that does not retain its origin and transformations may produce an elegant report, but it gives neither the farmer nor the researcher a path to verification or correction. Agricultural data are not free fuel for artificial intelligence. They are traces of fields, seasons, labour, decisions, and rights. A sensor reading may carry information about a plant's condition, a price record may reflect a farmer's position in the market, and a local designation may carry a community's memory. Technical value must therefore not be separated from the context that produced it or from the people who bear the consequences of its use. A mature system preserves raw data without treating them as sacred, makes transformations visible without confusing interpretation with the original, and distinguishes zero from missing, measurement from estimate, and observation from approved record. It standardises units, dates, and languages without erasing differences, and links entities through stable identifiers without replacing local names or dispossessing their owners. Quality is not a permanent medal attached to a dataset, but fitness tied to a question and its consequences. What is sufficient to alert a farmer to inspect a patch may not be sufficient to calculate a dose; what is suitable for a report may not be suitable for training a model; and what is shown to a reviewer should not operate a pump. The closer data come to action, the stronger the duty to verify and review them and to retain the ability to stop. Provenance gives the system its ethical and technical memory. It enables a trace back from a recommendation to the model, from the model to its features, from the features to the records, and from the records to the file, device, and field. It reveals what was affected when an error was found, why a result changed, who authorised a transformation, and what must be withdrawn or rebuilt. Yet traceability is incomplete without rights that can actually be exercised. Saying ‘the data belong to the farmer’ is not enough if farmers cannot download them in a usable form, correct them, see who accessed them, refuse their use for model training, or leave the platform with their history and relationships intact. Portability is not a button, and exit is not merely a contractual promise; both are capabilities to be tested before they are needed in a crisis. In this context, cybersecurity is agricultural safety. When a platform controls water, ventilation, or cooling, tampering with a digital command becomes an act in the physical world. Control networks must therefore be segmented, least privilege applied, backups and restoration tested, the safe mode during an outage understood, and human operators kept able to stop and intervene without losing essential functionality. Local knowledge, meanwhile, must not be treated as raw material for a system to extract and re-present without its owners. It is preserved in its language and context, with its attribution; linked to reference concepts without being erased; and governed by rules on who may use it and what benefit returns to those who supplied it. Standardisation that erases difference does not create shared knowledge; it creates an orderly database at reality's expense. A good system can be summarised in four capabilities: remembering the original, explaining the transformation, constraining use, and enabling exit. Together, they allow artificial intelligence to support decisions within intelligible limits. Without them, an algorithm may accelerate collection, linking, and prediction while also accelerating error and deepening vendor lock-in. Ultimately, an agricultural data platform is measured neither by the number of rows it can hold nor by the number of models it can run, but by its ability to answer when a single number is asked: Where did you come from? What changed in you? What are you permitted to do? Who can correct, withdraw, or carry you elsewhere? When the answers are clear, data become a foundation for knowledge and sovereignty. When they are absent, the number remains but trust disappears.
Evidence Notes
[SRC035] supports the continuum perspective from sensing, data, and context to decision and action, and is used here to frame the importance of boundaries between system phases, not to prove every architectural detail presented in the chapter. [SRC032] supports discussion of farmers' rights, access, trust, interoperability, and portability. [SRC030] outlines the risks of cybersecurity in smart agriculture and the expansion of the attack surface as connectivity expands. [SRC033] guides the view of risk, audit and controls as responsibilities that span the lifecycle of a system. The observation-record design, quality cards, transformation log, exit tests, and platform checklist are architectural and editorial constructs intended to turn the principles of provenance and governance into practical questions and procedures. These constructs are not a complete legal standard and do not replace the regulatory and contractual requirements specific to each country, sector, and use.