Skip to book content
Menu
International Office · Istanbul, Türkiye dr.alaa@aladdin.my.id +90 541 514 37 21

Aladdin Interactive Book

Chapter Three: The Evidence Framework for Agricultural AI

0%
Book knowledge toolsSearch and discuss Part ISearch this language edition or ask Chat V2 a question grounded in the approved book passages.
Discuss with Chat V2Book-only mode is the default. Every supported answer must cite an actual indexed passage.From this book only

Chapter Three: The Evidence Framework for Agricultural AI

The Chapter's Message: From an Abundance of Sources to Trustworthy Knowledge

The strength of a scholarly book does not lie in the number of references accumulated in its footnotes, but in its ability to determine what each reference can actually prove and where its epistemic authority ends. A source derives its value not from its mere presence, but from its fitness for the claim it is invoked to support. This issue becomes more complex in a rapidly evolving field such as agricultural AI: the writer may have before them a peer-reviewed scientific paper, an official report, a product page, a code repository, a computer simulation, a news report, and testimony describing a field experience. Each of these sources may contain accurate information, but they do not all carry the same degree of reliability or serve the same inferential function. A scientific paper may measure a model's performance under specific data and experimental conditions, but it does not automatically prove its success in every crop or region. A regulatory authority is the competent body for registering a pesticide or determining whether its use is authorized, but it is not the body that measures an algorithm's accuracy. A product page documents what the supplier claims about its system on a given date, but it does not substitute for an independent evaluation of its effectiveness. As for a code repository, it may establish the existence of a component in the code at the time of inspection without establishing that the component operates in a live service, or that it produces tangible agricultural impact in the field. The task of this chapter follows from that distinction: to explain how the book moved from a diversity of sources to disciplined, traceable, reviewable judgments. It does not present an abstract catalogue of research methods; rather, it reveals the practical framework used to distinguish a documented fact from a context-bound result, a description attributed to its source, a figure not yet fully verified, or an editorial proposal that does not claim to be an experimental finding. The aim is not to weaken the argument through an accumulation of caveats, nor to surround knowledge with a permanent fence of doubt; it is to build trust grounded in a clear understanding of its basis and limits. Scientific certainty is not produced by a categorical tone, but by a pathway the reader can follow: where did the information come from? What kind of source is it? What does this source actually prove? And what are we not entitled to infer from it?

1. The Architecture of the Evidence Base: How Was the Book's Knowledge Foundation Built?

By the "multilingual source corpus", the book means the total body of materials gathered and organized during the stages of its preparation, on which its chapters and analyses were then built. This corpus was not confined to a single type of document; rather, it included primary studies, scientific reviews, official reports, working papers, institutional documents, product descriptions, technical materials, and implementation guides, alongside auxiliary materials used to discover names and topics and to open lines of inquiry. These materials came in multiple languages, because agricultural innovation is not published entirely in English, and because some applications, policies, and local terminologies do not disclose their full meaning except within their own language and context. Yet linguistic plurality does not mean counting a translation and the original as two independent sources, nor does it grant a document a higher scholarly rank merely because it appears in more than one version. For that reason, translations were linked to their originals, duplicate or derivative versions were identified, and each document was assigned its precise place and function in the construction of the argument. The work began with the inventory and organization of the source material, but it did not stop at gathering files. It moved on to building separate editorial registers for sources, claims, figures, and case studies. The difference between the two stages is fundamental: collecting a document means only that material has become available for inspection, whereas accepting a claim derived from it requires another level of verification. At that point, the questions become: is this source suitable for substantiating this specific claim? Can the figure be traced to its original location? And did the conditions and limits of the study remain visible when the result was transferred into the book? A separate register was also created for the evidence associated with the case study Aladdin Agri AI, to prevent the mingling of internal implementation evidence with the general scientific references available to the reader. The existence of code, an interface, or an integration pathway may establish a specific implementation state in a known version, but it does not thereby become an independent study of effectiveness, nor evidence of agricultural impact in the field. This separation protects the value of both types of evidence: technical evidence is not made to bear more than it can prove, while the reader can still see precisely what it does establish. When the files and their paths were reorganized, references were not left hostage to obsolete locations or mutable names; instead, the baseline record of file paths and digital fingerprints was re-established. The purpose was not purely technical, but editorial and epistemic as well: to ensure that every important claim remained linked to the material on which it was based, and that moving files or changing their names would not break the chain of verification. Why is this work described as an evidence-informed structured narrative review? This book takes the form of an evidence-informed structured narrative review: it gathers heterogeneous materials, classifies them, compares them, and then builds from them a coherent interpretation of the field, while making visible the type of each piece of evidence and the limits of what it permits one to infer. This method was chosen because it suits the broad question the book poses: how does artificial intelligence move from the laboratory and the product into agricultural decision-making, the economy, labour, governance, and the field? A systematic review, in the precise scientific sense, serves a different purpose and rests on a more specific methodological framework. It usually begins with a narrower question, a pre-specified search protocol, defined databases, reusable search strings, clear inclusion and exclusion criteria, and a complete record of search results, deduplication, and screening stages. These elements enable another researcher to repeat the process and compare the findings with those of the review's authors. The difference in classification does not mean that one type is rigorous and the other less so; it means that each has a different epistemic function. A systematic review is well suited to answering a specific question within a repeatable protocol, whereas a structured narrative review can connect domains not united by a single experimental design: algorithms, agriculture, economics, food safety, labour, data rights, institutional structure, and regional differences. The strength of this book lies in making those connections rigorous without effacing the differences among types of evidence or presenting them as equal in nature and evidential force. Inherited counts from earlier drafts: when precision matters more than the appeal of the number Some drafts of an earlier edition of this book stated that more than 150 documents had been examined across multiple databases and sources. Yet this edition did not retain that number merely because it appeared in the earlier version: it was not accompanied by a record complete enough to reproduce the search strings, dates, results, screening, exclusions, and deduplication required to treat it as a final methodological count. The easier choice would have been to carry the count into the new edition; it is a large, clear, and attractive number. The scholarly choice, however, was to separate useful knowledge that could be verified from a numerical claim whose documentary pathway remained incomplete. Therefore, the historical number was not used as evidence of the comprehensiveness of the review, and the materials and claims were reassessed in accordance with this edition's records and stated limits. This does not mean that the knowledge that passed through earlier drafts was wasted or nullified. Names, ideas, and research pathways that could be traced were retained, and claims were re-examined before being adopted. As for any number or result that could not be traced to an appropriate source, it was not presented to the reader as a fully established fact. Thus editorial review became a tool for refining knowledge, not erasing it; a means of raising the book's reliability, not diminishing it. A published systematic review such as [SRC004] may contribute to mapping part of the field and may offer an example of a documented search protocol within its own specific question. Yet its systematic character belongs to the work carried out by its authors, and does not automatically pass to this book merely because it cites it. Similarly, a broad collection of documents does not become a systematic review merely by virtue of the number of its files or the orderliness of its chapters. Precision begins with naming the work for what it has actually accomplished, and then holding it rigorously accountable to the demands of that name. In this sense, method is not an apology for the book, but a statement of its strength: the book chose to build a broad, traceable argument, to declare its nature clearly, and to give every source its due place without elevating it beyond its actual evidential force. Methodological conclusion: trust is not built from the sheer number of documents, but from the clarity of the path that links every claim to its source, and specifies precisely what that source proves and what lies beyond its limits.

2. The Unit of Verification Is the Claim, Not the Paragraph

A paragraph on the page may appear to be a single coherent unit, yet it may contain several claims that differ radically in nature and in the type of evidence required to establish each one. A sentence that establishes the existence of a system is not on the same level as a sentence that describes one of its properties, and neither is sufficient to prove its accuracy or its impact on farmers' lives. Let us consider a simplified example of a paragraph dealing with a digital agricultural product. The paragraph may include four different levels of claim:

  • A claim of identity or existence: "There is a platform or tool bearing this name."
  • A claim of property: "Its provider states that it uses computer vision to detect symptoms of disease."
  • A claim of performance: "It achieved 95% accuracy on a specific task."
  • A claim of impact: "It reduced farmers' losses or improved their income under actual operating conditions."

A supplier's page may document that the product is offered under this name, and that the company describes it as using computer vision. But in that case it documents only what the supplier says about its product at a given date. The page alone does not provide independent verification that the feature works as described, or that the system achieved a given level of accuracy, or that its use in fact led to reduced losses. As for the performance figure, it requires a source that specifies the data on which the model was tested, how the data were split, the definition of the metric used, the baseline against which the result was compared, and the limits of generalizing it beyond the study setting. Even if it is established that the model achieved an accuracy of 95% in a particular test, that result does not automatically prove economic or agricultural impact. Proving a reduction in losses requires a design that measures what happened in the field, compares the result with what would have happened in the absence of the system, and takes into account the costs, risks, and other factors that might explain the improvement. Three questions govern the examination of every claim For this reason, the book's method treats the verifiable claim as the basic unit of examination. This means breaking down the composite statement into clear units, and then putting three questions to each unit:

  1. What type of claim is it?
  2. Is it a claim of identity, property, performance, or impact?

  3. What is the appropriate source for proving it?
  4. Does it require an official document, a performance study, a field evaluation, or technical material?

  5. What are the limits of the formulation permitted by the source?
  6. What can be said with confidence, and what would count as a generalization that exceeds the evidence?

After examination, the claims are not all grouped together in a single category, but classified according to clear grades:

  • Verified: supported by suitable material, and its chain of proof can be traced.
  • Verified within limits: established in a specific study or context, and not to be generalized beyond it.
  • Attributed description: expresses what an entity or company says about its product, not an independent evaluation of it.
  • Awaiting verification: there is an initial indication or reference, but the chain of proof is incomplete.
  • Unusable: no matching source could be identified, or the source was found not to substantiate the claim attributed to it.

So if the product name and its field of use are established, while its performance level is not, the product need not be removed from the map of the field; but at the same time, the figure may not be published as though it were a complete fact. The name and the task description are retained within the limits that could be documented, while the performance figure is temporarily isolated from the main text until an appropriate source for it appears. Scientific quarantine: protecting the conclusion before the evidence is complete This is what is meant by scientific quarantine in this book. It is not a judgment that the idea is false, nor an accusation against the source or the product, but an editorial decision that prevents an incomplete claim from operating within the text as a fact. A claim may be reopened if a matching primary source or sufficient independent verification emerges, and it may remain outside the conclusion if the conditions for establishing it are not met. In this way, scientific quarantine becomes a mechanism for protecting knowledge from haste, not a tool for excluding ideas. On this basis, rigorous control of the evidence does not impoverish the content, just as the desire to present a wide-ranging book does not become a justification for conflating degrees of certainty. The book can retain a rich map of applications, products, and research while, at the same time, drawing a clear line between:

  • what we know on the basis of appropriate evidence;
  • what we know within specified conditions and contexts;
  • what the party says about itself;
  • and what still awaits sufficient substantiation.

Scientific rigour does not make knowledge less expansive; rather, it makes its boundaries clearer. When the reader knows where a statement came from, what supports it, and where its scope ends, caution no longer appears as a sign of weakness, but becomes the foundation that gives confidence its meaning.

3. From Model Accuracy to the Quality of Agricultural Decision-Making

An artificial intelligence model may achieve high results in tests, especially when it is evaluated on clean, complete data resembling the data on which it was trained. Yet this technical success alone does not guarantee that it will change an agricultural decision in practice. The model may be impossible to run on the phone or device available in the field, may require a constant internet connection, may issue its alert only after the opportunity to save the crop has passed, or may offer a confident recommendation even though the image is unclear or the data are incomplete. In such cases, the numerical result remains impressive, but its practical value is limited. Conversely, another model may record slightly lower test performance, yet be lighter to run, better able to operate under field conditions, and more candid about the limits of its knowledge. If the data are insufficient, it refrains from issuing a categorical recommendation and asks for an additional image or measurement, or refers the case to an agronomist. If it detects a risk, it issues an early warning that gives the farmer sufficient time for inspection and action before the intervention window closes; that is, the period during which treatment or prevention remains possible and effective. This model may therefore be more useful, even if it does not have the highest score in the test results. Consider a practical example: a model for diagnosing a plant disease may identify the infection with high accuracy, yet detect it only after the symptoms become clear and the disease has already spread over a wide area. In practice, it may be outperformed by a slightly less accurate model that captures the early signs, states its level of confidence, and distinguishes a case that can be monitored from one that requires urgent human examination. The first describes the problem well after it has occurred; the second helps people decide while the decision can still affect the outcome. This does not mean belittling the importance of accuracy or accepting a weak model on the pretext that it is easy to operate. Technical performance remains an essential condition, especially in decisions that may affect the safety of crops, animals, or people. But agricultural value is not produced by a single figure; it is produced by an integrated system that combines acceptable and substantiated accuracy, context-appropriate data, reliable operation, timely warning, an understandable and actionable recommendation, and responsible abstention when the evidence is insufficient. The best model, then, is not always the one that knows more inside the laboratory, but the one that helps people make a sounder decision at the place and time when that decision can still change the outcome. For this reason, five levels of success should be distinguished:

Level of evaluationThe farmer's direct questionWhat does success look like in practice?What does this success alone not prove?
1. Model accuracy in the testWas the model able to identify the disease, predict the yield, or estimate the irrigation requirement correctly in the data on which it was tested?For example, that it distinguishes between healthy and infected plants with few errors, or that its yield predictions are close to the actual values in the test.It does not prove that it will achieve the same performance when used on another farm, or with a different variety, or in a different season, weather, and lighting.
2. Its suitability to farm conditionsHas the model been tried on new data resembling the conditions of my farm, rather than only on the data from which it learned?That it is tested on farms, regions, or seasons not included in training, or that it undergoes a local trial on the target crop and environment and maintains an acceptable level of performance.It does not prove that the application, device, or sensor will continue to operate steadily every day, or under weak internet, power outages, and data scarcity.
3. System performance in operationDo the data, alerts, and recommendations arrive complete and at the time I need them?That the sensors and applications operate stably, and that a disease or irrigation alert arrives before the time for intervention has passed, with a safe fallback when the connection or service fails.It does not prove that the farmer understood the alert, trusted it, or was able to carry out the recommendation with the time, resources, and equipment available.
4. Improvement of the agricultural decisionDid the recommendation help the farmer take a better action than the usual one?That the alert prompts an early inspection of a suspicious patch, or adjustment of the timing and amount of irrigation, or avoidance of an unnecessary treatment, or referral of a sensitive case to a specialist.It does not by itself prove that the crop, profit, or resource consumption improved; the decision may be sound, yet its effect may be hindered by other factors such as weather, lack of inputs, or delayed execution.
5. Actual impact on the farmAfter using the system, did the outcome that actually matters to the farmer improve?That losses or total production cost decline, or crop quality or yield stability improves, or water and energy consumption falls, or safety risks decline, compared with a clear baseline.It does not prove that the effect will automatically be replicated on every farm or in every season; the result may vary with the crop, climate, holding size, prices, and method of use.

These levels are not different names for the same success, but links in a single chain. A model may be accurate yet unsuited to the local environment; it may suit that environment but the service may fail; the service may function but the recommendation may arrive too late or remain vague; and the farmer may make a better decision without its benefit becoming visible in an exceptional season. Therefore, no judgment on the technology is complete unless the whole path is followed: from the correctness of the prediction, to its suitability for the farm, to its timely arrival, then to the decision it changed, and finally to the effect it had in the field and in the farmer's accounts. The evidence chain can be represented as follows: Valid data     ↓ An appropriate model     ↓ Independent verification     ↓ Reliable operation     ↓ An understandable recommendation     ↓ A feasible and timely action     ↓ An agricultural outcome     ↓ Economic, environmental, and social impact A claim is not entitled to leap over a missing link. If the available evidence concerns model performance, the conclusion is framed within those limits. If the evidence is a short field trial, it is not presented as stable operation across seasons. And if yield changes, the change is not attributed to the system before considering the weather, the area, the inputs, and the baseline.

4. Different Metrics for Different Tasks: What Does the Reported Number Mean?

Studies and product pages are full of phrases such as "high accuracy", "advanced performance", and "results that outperform previous models". Yet a number, however dazzling it may seem, does not explain itself. Before we ask, How high was the accuracy? we should ask an earlier question: What task was being measured, and what kind of error can arise in it? Artificial intelligence may be used to answer entirely different questions:

  • Is the plant infected or healthy?
  • Where is the fruit or the insect within the image?
  • What is the area of the affected part of the leaf?
  • What is the expected yield at the end of the season?
  • Can the robot pick the fruit without damaging it?
  • Did the irrigation system actually reduce water consumption?

These tasks cannot all be measured by the same yardstick. Nor is it permissible to place their numbers in a single table and rank applications from highest to lowest as though they were competing in a single test. Each metric illuminates one aspect of performance, while leaving other aspects in the shadows. The reader does not need to memorize the mathematical equations behind these metrics. What matters is to understand three things: what question the metric answers, what error it may conceal, and how it relates to the agricultural decision. First: when the task is to classify a condition — is the plant diseased or healthy? Let us suppose that an application receives an image of a plant leaf and classifies it into one of two states: "diseased" or "healthy". This may seem simple, but the model can produce four different outcomes:

Actual conditionWhat the model saidWhat does the result mean for the farmer?
The plant is diseasedIt said it was diseasedA correct detection that may lead to inspection or early intervention
The plant is diseasedIt said it was healthyAn infection the model missed, and the disease may continue to spread
The plant is healthyIt said it was healthyA correct exclusion, which may spare the farmer an unnecessary inspection or treatment
The plant is healthyIt said it was diseasedA false alarm that may consume time or lead to an unnecessary inspection or treatment

These four outcomes are the basis from which most classification metrics are derived. The differences among the metrics are not a statistical luxury; each of them looks at a different aspect of error. Overall accuracy: how many answers did the model get right? Overall accuracy is the proportion of all predictions that are correct, whether they are diseased cases correctly detected or healthy cases correctly excluded. If the model examines one hundred images and gets ninety of them right, its overall accuracy is 90%. This number seems clear, but it can become misleading when one class is far more numerous than the other. Suppose, for example, that we have one hundred images: ninety-five showing healthy plants and only five showing diseased plants. If a weak model declares that all the images are healthy, it will be correct in ninety-five cases, yielding an overall accuracy of 95%. Yet it has not detected a single one of the five infections. So overall accuracy answers the question: How often was the model correct overall? But it does not by itself answer the more important question in this example: How many real infections was the model able to detect? Recall/Sensitivity, or the ability to detect diseased cases Recall, also known as sensitivity, measures the proportion of all genuinely diseased cases that the model successfully detected. In the farmer's language, the question is: If the disease is actually present, how likely is the system to detect it rather than let it pass without an alert? Recall is especially important when missing a diseased case is highly dangerous—for example, with a rapidly spreading disease, an early sign of a health disorder in a herd, or a pest that can be contained if detected in time. High recall means that the model misses fewer diseased cases. But it does not necessarily mean that every alert it issues is correct; the system may increase recall by issuing many alerts, some true and some false. Specificity, or the ability to exclude healthy cases Specificity measures the proportion of healthy cases that the model correctly identified as healthy. The question here is: If the plant is healthy, what is the probability that the system will not misclassify it as diseased? Specificity matters when false alarms carry a high cost, such as sending teams for inspection, disrupting production operations, taking laboratory samples, or considering an unnecessary treatment. But high specificity alone does not guarantee that the model detects diseased cases. It may be very conservative in issuing alerts, thereby reducing the number of false alarms, but at the cost of missing real infections. Precision: how many positive alerts were correct? In Arabic, this metric is sometimes rendered as iḥkām ("exactness"), but that term may not be clear to readers. Its practical meaning is simpler than the label suggests: Out of all the cases the system declared to be diseased, how many were actually diseased? If the system issues ten alerts, and inspection shows that only six are correct, then its precision is not perfect: four alerts were false. This metric matters when following up an alert is costly or may lead to a sensitive intervention. If every alert calls for laboratory analysis, a specialist visit, or the stoppage of part of the production line, it is not enough for the system to capture most diseased cases; it must also avoid drowning the user in false alarms. Here the difference between precision and recall becomes clear:

  • Recall asks: out of the cases that were actually diseased, how many did the model detect?
  • Precision asks: out of the cases the model declared diseased, how many alerts were correct?

A model may be strong on one side and weak on the other; therefore, the choice of metric must be tied to the real-world consequences of error in the agricultural context. F1 score: does the model balance detecting infection with the correctness of its alerts? The F1 score combines recall and precision into a single value. It rises when the model succeeds in two things at once:

  • Detecting a good proportion of diseased cases.
  • Avoiding the issuance of a large number of false alerts. It may be viewed as a concise score for balance between not missing infection and not overstating warnings about it. If the model detects every infection but issues many false alerts, its F1 score will not be as high as it should be. The same is true if its few alerts are correct but it misses many infections. Yet combining the two dimensions in a single number comes at a price: it may conceal the fact that one type of error is more dangerous than the other. In a rapidly spreading epidemic disease, missing a single infection may be more harmful than sending several healthy cases for inspection. In a high-cost or high-risk intervention, however, the false alarm itself may become a serious problem. The F1 score therefore does not say that the model is automatically "safe" or "suitable"; it says only that it has achieved a degree of mathematical balance between two kinds of performance. The balance required in agriculture is determined by the consequences of error and the nature of the decision.

The confusion matrix: a table showing where the model was right and where it was wrong Despite its confusing name, the confusion matrix is not a mathematical riddle, nor is it a single metric like overall accuracy. It is simply a table that gathers the four outcomes above:

  • Diseased cases correctly detected.
  • Diseased cases missed by the model.
  • Healthy cases correctly excluded.
  • Healthy cases incorrectly classified as diseased. The value of the confusion matrix lies in the fact that it does not hide errors behind an aggregate number. Two models may achieve the same overall accuracy, yet one misses dangerous diseases while the other issues many false alerts. The aggregate number may make them seem similar, but the table of errors reveals the difference that matters to the farmer. Hence the simpler name that may accompany the term is:

Confusion matrix: a table that shows the types of correct and incorrect answers. How do we choose the appropriate metric in diagnosis? There is no single best metric in all cases. The choice depends on the agricultural question and on the consequence of each type of error.

  • If missing the disease could lead to extensive spread or to losses that are hard to reverse, Recall becomes a high priority.
  • If responding to every alert is costly or may lead to unnecessary intervention, Precision and Specificity become more important.
  • If the classes are imbalanced, overall accuracy alone is not sufficient.
  • If we want a comparative figure that summarizes the balance between detecting cases and the correctness of alerts, we can use F1, while referring back to the details of the errors.
  • If we want to know exactly where the model went wrong, we turn to the confusion matrix.

The application must not move from initial recognition to a sensitive treatment recommendation on the basis of this evaluation alone. A model's success in distinguishing a suspicious image does not, by itself, establish the correctness of the causal diagnosis or the suitability of the proposed intervention. There may still be a need for field inspection, laboratory analysis, specialist review, and reference to the official label and local regulations. Second: when the task is to determine the location of an object or draw its boundaries Classification usually answers a question such as: "Is there an insect in the image?" But some applications require more precise information:

  • Where exactly is the insect located?
  • How many fruits appeared in the image?
  • What is the position of the weed between the crop rows?
  • What area of the leaf is affected?
  • Where does the healthy part end and the damaged tissue begin? Here we move from classification to two more detailed visual tasks:
  1. Detection: the system identifies the object's location within a box or approximate area, such as drawing a box around each fruit.
  2. Segmentation: the system draws the precise boundaries of the object or region, such as identifying the area of a disease spot on the leaf in greater detail. Because the question is no longer merely "Is the object present?", we need measures that compare the object's location and boundaries with the correct location determined by experts in the reference images.

Intersection over Union (IoU): to what extent did the predicted region match the correct region? Suppose an expert drew the correct boundaries of a disease spot on a leaf, and then the model drew a region it predicted represented the same spot. The Intersection over Union measure compares the two areas:

  • the shared part between what the expert identified and what the model identified.
  • the total area covered by the two regions together. The greater the true overlap and the smaller the extra or missing part, the higher the value of the measure. It can be understood visually as follows:
  • complete overlap between the two regions means ideal performance in determining the boundaries.
  • partial overlap means that the model got one part right and another part wrong.
  • no overlap means that the model identified an entirely different location. But a high value on this measure does not answer the question of whether the region was identified with sufficient precision to carry out a particular agricultural task. It may be suitable for estimating the extent of the affected area, yet insufficient for guiding a robotic arm that requires higher spatial precision.

Dice coefficient: what degree of similarity is there between the two areas? The Dice coefficient also measures the degree of similarity between the region identified by the model and the reference region identified by the expert. It is close in concept to Intersection over Union, even if the method of calculation differs. The reader does not need to memorize the mathematical difference between the two measures. What matters is that both attempt to answer the question: To what extent did the model draw the correct region, without leaving out important parts or adding areas that do not belong to it? The Dice coefficient appears frequently in image-segmentation tasks, including the identification of tissues, leaves, fruits, or affected areas. But it remains a measure of visual similarity, not complete evidence of the success of the agricultural decision that will be built on this segmentation. Mean average precision (mAP): how did the system perform in finding objects and determining their locations? mAP is often used when evaluating object-detection systems in images, such as detecting fruits, insects, weeds, or animals. It is a composite measure that summarizes detection quality across multiple classes or evaluation conditions. Rather than entering into its computational formula, its function can be understood as follows: Did the system find the required objects, avoid detecting objects that are not there, and determine their locations with an acceptable degree of overlap with the correct locations? The mAP value should not be read as a direct "agricultural success rate". If a model achieves a high figure in detecting fruits within a set of images, that does not mean the robot will successfully pick the same proportion of fruits. The vision system may see the fruit well, then the arm may fail to reach it, or grasp it too firmly, or take an uneconomical amount of time, or be unable to work under the lighting, dust, and branch movement in the field. In a robotic system, vision is only a first link, followed by other links: seeing the target ← determining its position ← planning the movement ← reaching it ← executing the action ← avoiding damage ← completing the task within acceptable time and cost. For this reason, vision accuracy must be distinguished from the success of the full agricultural operation. Third: when the task is to predict a number — yield, price, or demand In some applications, the model does not classify the plant as "infected" or "healthy", nor does it look for a fruit within an image; rather, it predicts a numerical value, such as:

  • the expected yield quantity.
  • the crop's water requirement.
  • the expected price over a given period.
  • the volume of demand for a product.
  • the remaining storage life.
  • the amount of feed or energy required. In these cases, we need to know the size of the difference between the number predicted by the model and the number that actually occurred. There are several ways of calculating this difference, because each method reveals a different aspect.

Mean absolute error (MAE): how much does the model err on average, in a unit the farmer understands? If the model predicted a yield of 5.5 tonnes per hectare, while the actual yield was 6 tonnes, then the error in this case is half a tonne per hectare. Mean Absolute Error adds up the magnitudes of these errors across all cases, without regard to whether the model overestimated or underestimated, and then calculates their average. Its main advantage is that it is expressed in the unit of the variable itself:

  • tonnes per hectare when predicting yield.
  • lira per kilogram when predicting price.
  • cubic metres when predicting water consumption.
  • days when predicting storage life. For that reason, it is one of the measures most readily translated into a practical question:

If I rely on this model, how far, on average, will its prediction be from reality? But the average may conceal a few cases in which the model made very large errors; this is where another metric is needed, one that gives those errors greater weight. Root mean squared error (RMSE): are there large errors that the average should not hide? RMSE also measures prediction error, but it penalizes large errors more than small ones. If most of the model's predictions are close to reality, but it fails severely in some seasons or regions, the effect of those failures will appear more clearly in this measure. Its practical idea is: We do not want only to know the average error; we also want large errors not to pass as though they were ordinary cases. This becomes more important when a large jump in error affects the decision. A limited error in predicting yield may be manageable, but a large deviation may lead to an unsuitable sales contract, insufficient storage capacity, or a mistaken estimate of transport and financing. As with MAE, RMSE is expressed in the unit of the variable itself, but its value is affected to a greater degree by cases with large errors. Mean absolute percentage error (MAPE): what percentage error is there relative to the actual value? MAPE converts the error into a percentage, which makes it attractive and easy to present. The phrase "the average error reached such-and-such percent" seems more intelligible than a figure in a unit that may differ across crops and regions. But this measure may become misleading when the actual value is very small or close to zero. A simple numerical error may, when divided by a small value, turn into a huge percentage that does not reflect the practical significance of the case. Therefore, before relying on it, we should ask:

  • Can the actual values approach zero?
  • Is percentage really the most suitable way to understand the error?
  • Was the magnitude of the error in the original unit also reported?
  • Does the percentage conceal important differences between crops, regions, or seasons? Ease of reading the percentage does not mean that it is always suitable.

The coefficient of determination (R²): how much of the variation in the outcomes was the model able to explain? The coefficient of determination measures the extent to which the model, within the data and evaluation method used, is able to explain variation in the observed values. Suppose yield varies across fields and seasons because of multiple factors. The model attempts to explain part of this variation on the basis of the data fed into it, such as weather, soil, cultivar, and farming practices. The coefficient of determination shows how much of the variance the model was able to represent in comparison with a statistical baseline within the testing framework. But this metric is often misinterpreted. If a study reports that R² reached 0.92 in the evaluation setting it used [SRC006], it is not valid to turn that into the statement: "The system is 92% accurate." The two statements do not mean the same thing. An R² value does not directly tell us:

  • By how many tonnes or kilograms the model was wrong.
  • Whether the errors were economically acceptable.
  • Whether it succeeded in a drought season or a heatwave.
  • Whether it maintains its performance in another region.
  • Whether the prediction arrived in time to allow the decision to be changed.
  • Whether it was better than a historical average or a simpler, less costly rule. For this reason, it is preferable to read the coefficient of determination alongside a metric that shows the magnitude of error in an intelligible unit, such as MAE or RMSE, and alongside a clear description of the data, place, season, and testing method.

A practical summary of the metrics

Type of taskMetricThe simple question it answersWhat to be cautious about
Classifying a plant or animalOverall accuracyHow many cases did the model classify correctly overall?It may appear high if healthy cases vastly outnumber diseased ones
Disease detectionSensitivityOf the actually diseased cases, how many did it detect?It may rise at the cost of more false alarms
Identifying healthy casesSpecificityOf the healthy cases, how many did it identify correctly?On its own, it does not prove the model's ability to detect disease
Assessing the correctness of alertsPrecisionOf all disease alerts, how many were correct?It may be high while missing a number of infections
Balancing detection and alert precisionF1Did the model maintain a balance between detecting disease and reducing false alerts?It conceals differences in the cost of the two types of error
Analysing classification errorsConfusion matrixWhere did the model get it right, where did it miss the disease, and where did it raise a false alarm?It is not a final judgment on agricultural usefulness
Locating an object in an imagemAPDid the system find the targets and locate them acceptably across the tests?It does not prove the success of the robot or the agricultural operation as a whole
Comparing two regions in an imageIoUHow much did the region identified by the model overlap with the correct region?By itself, it does not determine whether the accuracy is sufficient for the intended use
Drawing the boundaries of an affected areaDiceTo what extent did the predicted area resemble the reference area?It measures visual similarity, not the downstream consequences of the decision based on it
Predicting yield or priceMAEHow large is the prediction error on average, in an intelligible unit?It may conceal some large errors
Revealing large errorsRMSEAre there large deviations that should be given greater weight?It may be heavily affected by a limited number of extreme cases
Expressing error as a percentageMAPEWhat proportion of the actual value does the error represent on average?It becomes unstable or misleading near zero
Explaining variation in valuesHow much of the variation was the model able to explain within this evaluation?It is not a percentage of accuracy, and it does not show the magnitude of error in an agricultural unit

The right number is not enough if it is unfit for the decision After understanding the metric, the more important question remains: did the evaluation measure the attribute the farmer actually needs? A yield prediction may be close to reality, yet arrive after the time for buying inputs or arranging storage and marketing contracts has passed. A vision model may detect a pest efficiently, yet fail to distinguish between a case that calls for monitoring and one that requires urgent intervention. An automated system may determine the fruit's position accurately, yet its arm may be unable to reach it without damaging the branches. An irrigation system may reduce the amount of water used per kilogram of yield while not reducing total withdrawals from the water source. For this reason, after reading any technical metric, the following questions should be asked:

  1. What task exactly did the number measure?
  2. Did it measure image classification, object localization, quantity prediction, or the success of a complete agricultural operation?

  3. On what data and in what environment was the test conducted?
  4. And do they represent the crop, region, season, and equipment in which the system will be used?

  5. What kinds of errors does the aggregate number conceal?
  6. And is the more serious issue missing the case, raising a false alarm, or delaying the recommendation?

  7. Did the output arrive before the intervention window closed?
  8. A correct prediction after the time for irrigation, control, or harvest has passed is belated knowledge, not a useful decision.

  9. Did the system perform better than a simpler alternative?
  10. Such as routine inspection, a clear agronomic rule, a historical average, or a lower-cost advisory service.

  11. Was the farmer able to understand the recommendation and carry it out?
  12. Technical accuracy does not turn into benefit if the recommendation is vague, requires unavailable resources, or ignores local constraints.

  13. Did the benefit remain after accounting for cost and risk?
  14. Including the price of devices, connectivity, maintenance, subscription, data entry, training, and potential errors.

In sum, metrics are not arcane formulae the reader is required to memorize, nor medals that grant technology automatic legitimacy. They are tools for answering specific questions. A discerning reader is not dazzled by a high number before asking what was measured, what was omitted, who bears the cost of error, and whether that performance could bridge the distance between the test screen and a useful decision in the field.

5. The Maturity Ladder: From a Claimed Feature to Demonstrated Agricultural Impact

Technical writing sometimes reduces the state of a product to two words: "present" or "absent". But between the idea and agricultural impact lie many levels, and each level has evidence appropriate to it. This book proposes an editorial ladder for understanding maturity. It is not a formal regulatory standard, but a tool that helps the reader prevent a leap from the existence of a feature to a claim of impact.

LevelWhat does the evidence prove?What does it not yet prove?
1. Stated purposeThe organization says the system was designed for a specific taskThat the feature is implemented or works
2. Implemented componentThere is code, an interface, or a prototype that performs part of the taskThat it works outside the technical demonstration or the laboratory
3. Internal performanceThe model was evaluated on the developer's or the study's dataThat its performance will transfer to independent data
4. External validationTested on data or by a party outside the model-building processThat it will endure in day-to-day agricultural operation
5. Field trialUsed on a real farm or within a real value chain under a known duration and stated conditionsThat it will continue to perform across seasons, locations, and users
6. Recurrent operationOperates within a service that has monitoring, support, recovery, and clear responsibilitiesThat it achieves a net economic or environmental benefit
7. Measured impactAn agricultural, economic, or safety outcome changed relative to an appropriate baselineThat the impact will recur in every context

It is not fair to demand of a research model commercial operational evidence it does not claim, just as it is not acceptable for a commercial product to use the result of a prototype model as proof of field impact. A system can be advanced on one level and lagging on another. Its software architecture may be mature, while the agricultural evidence remains preliminary. A field trial may succeed, while data rights or the maintenance plan remain immature. Therefore, readiness is not reducible to a single score when its dimensions differ. Three questions reveal the illegitimate leap When reading any description of an application, ask:

  1. What is the highest level the evidence has actually established?
  2. To what level does the marketing statement leap?
  3. What is the missing link between the two?

If the source establishes the existence of a component while the statement speaks of reducing losses, what is required is not improved wording, but evidence spanning the distance that has not yet been established.

6. Regional Transferability Limits: Moving Models Between Agricultural Environments

Restricting the result of an agricultural AI model to the testing conditions in which it was produced may seem excessively cautious, but the importance of that caution becomes clear when one asks the practical question: Will the algorithm maintain its performance when it leaves the training and testing environment and is used in a field that differs in climate, soil, crops, and farming practices? The model does not learn "agriculture" in the abstract; rather, it learns statistical relationships within data that came from specific soils, climates, varieties, devices, and methods of measurement. When one of these elements changes, the model does not necessarily fail, but its previous performance loses its validity as a direct promise and reverts to a hypothesis that requires fresh verification.

Dimension of differenceWhat changed?Potential effectWhat is required before reliance
Soil and climateMoist temperate soil ← dry or saline soilSpectral reflectance and response curves changeLocal samples, calibration, and ground verification
Holding sizeLarge homogeneous farm ← small mixed holdingThe assumption of homogeneity weakensRetraining or narrowing the scope of use
Connection qualityStable network ← intermittent connectionCloud inference fails and data become staleSafe local mode and a record of the last synchronization
Variety and genetic typeLimited commercial varieties ← local or heirloom varietiesSymptoms or appearances fall outside the training classesAn "unknown" class and an escalation pathway
Camera and sensorDevice with a specified standard ← cheaper or older deviceShift in lighting, resolution, and calibrationTesting on the actual device
Regulation and registrationOne jurisdiction ← another jurisdictionA recommendation is unregistered or locally prohibitedReview the label and consult the competent authority
Language and terminologyStandardized term ← local names for crop and pestMisunderstanding of the question before inferenceComprehension testing in the user's language
Season and cycleOne or two seasons ← multi-year variabilityPerformance untested in an exceptional yearVerification across seasons and drift monitoring
Method of workTrained specialist ← user with different experience or time constraintsThe quality of input and response changesUsability testing and appropriate training
Market and institutional structureAvailable intervention and established financing ← limited alternativesA correct recommendation that is not feasible to implementAnalysis of implementation capacity and cost

This table should not be read as a list of prohibitions, but as a testing plan. The wider the gap between the training environment and the environment of use, the greater the need for a limited local trial before widespread adoption. Local verification plan before adoption The plan may include:

  1. Selecting a sample that represents the target fields, varieties, and conditions.
  2. Testing on the devices and networks people will actually use.
  3. Comparing the system with current practice or with an appropriate simple rule.
  4. Examining performance in rare classes and unfamiliar cases.
  5. Providing an "unknown" class or an abstention mechanism.
  6. Defining the cases that require specialist review.
  7. Monitoring performance change over the course of the season.
  8. Recording outages, corrupted readings, and human interventions.
  9. Measuring the agricultural outcome, not the technical metric alone.
  10. Reassessing if the variety, device, data source, or regulatory environment changes.

Requiring local verification does not mean that the technology has no value outside its place of development; it means that its value must be demonstrated where people and the land will bear the consequences of the decision.

7. Error Is Not a Single Number: Who Pays for the Wrong Decision?

AI errors are not equal in meaning or consequence. Two systems may err at the same rate, yet one causes limited inconvenience that can be remedied, while the other opens the door to crop loss, resource waste, or risks to food, animals, or workers. It is therefore not enough to ask: How many times did the model get it wrong? We should also ask: In which direction did it err, when did it err, what decision followed, and who bore the cost? In the laboratory, an error is merely a case within a table. On the farm, it may become water used in the wrong place, a disease left undetected, a crop sprayed without need, a shipment whose cooling was delayed, or a worker who received an alert after the intervention window had closed. Risk is therefore measured not by the number of errors alone, but by the weight of their consequences. Two models may appear close in overall accuracy while differing radically in practical value. The first may issue excess alerts yet rarely miss a dangerous case; the second may be less bothersome, yet sometimes remain silent when silence is costly. Which is better? The aggregate number cannot answer on its own; the answer depends on the nature of the task, the cost of inspection, the danger of delay, the reversibility of the decision, and the human capacity to intervene. Error may arise in the model, or be produced by the surrounding system It is easy to attribute the whole error to the "algorithm", but the digital agricultural system is an interconnected chain: Sensor or image ← transmitted data ← model analyses them ← alert or recommendation ← user interprets it ← action is executed ← outcome appears in the field. The algorithm may be correct while the sensor reading is old or defective. The analysis may be correct, but the application displays it in the wrong unit of measurement. The appropriate recommendation may reach the farmer's phone after the intervention window has closed. The user may understand the alert as a definite order, though it was only a preliminary signal requiring inspection. Indeed, the error may occur after the model has completely exited the scene. The system may send a correct command to the irrigation device, but a fault in the connection causes the command to be repeated. The quantity may be recorded in litres per hectare, while another program interprets it in a different unit. The system may rely on a moisture reading from a sensor that has not been calibrated for some time, and thus build a logical inference on incorrect information. For this reason, three levels should be distinguished:

  • Model error: it erred in classification, estimation, or prediction.
  • Decision error: the information was insufficient, or it was interpreted in a way that did not suit the case.
  • System error: the data, software, hardware, modules, or execution paths failed, turning a correct or acceptable result into an erroneous action. This distinction is necessary because each level requires a different remedy. Improving the model does not fix a faulty sensor, adding more data does not remedy an ambiguous interface, and raising accuracy does not prevent an instruction from being executed twice because of a software fault.

False positive: an alert that a problem exists when it does not A false positive occurs when the system declares the presence of a disease, pest, fault, or risk when the condition is not actually present. Put more simply: the system raised an alert, but the subsequent inspection found no such problem. The consequence of this error may be limited. If, for example, the application suggests that the farmer visit a nearby part of the field to verify a suspected spot, and the inspection is quick, safe, and low-cost, the harm may amount to no more than some time and effort. In that case, some degree of over-alerting may be acceptable if it helps to avoid missing serious cases. But a false alert does not always remain simple. Its consequence becomes greater the closer it comes to automated execution or the higher the cost of responding to it. Examples include:

  • a visual model interprets the symptoms of a nutrient deficiency as a fungal disease, prompting the user to consider a treatment that does not address the real cause.
  • an irrigation system interprets an inaccurate reading as water stress and recommends irrigating soil that still contains sufficient moisture.
  • an animal-monitoring system repeatedly raises an alert about a non-existent health condition, consuming the time of the worker or veterinarian and leading to the unnecessary isolation of an animal.
  • a cold-chain system interprets an anomalous reading from a defective sensor as a real temperature rise, prompting an inspection, shutdown, or costly operational intervention.
  • a vision system detects a weed in a shadow or among plant residues, then sends its location to a removal machine or a spot-spraying unit. The consequence of the same alert therefore varies according to what follows. If it leads only to human inspection, it may be tolerable. But if it directly triggers a pump, changes the climate in a greenhouse, applies a chemical, or excludes a shipment, then the false alert becomes a potential source of economic, environmental, or safety-related harm. Nor is its damage confined to each individual incident. Repeated false alerts create what is called alert fatigue: the user becomes accustomed to hearing the alert and then begins to defer or ignore it. At that point, the system may lose credibility before the moment arrives when its alert is correct and important.

False negative: a real problem that the system did not see A false negative occurs when the problem is present, but the system does not detect it or classifies the case as normal. Put directly: there was a problem that warranted attention, but the system remained silent. This error may be more dangerous than a false alert, especially when losses grow over time. Examples include:

  • an early plant infection that the application failed to detect, so that it continued until it spread from a limited number of plants to a wider area.
  • a pest that did not appear clearly in the image or was not represented in the model's training data, so the system classified it as a normal condition.
  • the onset of water stress that the system failed to detect, so the plant passed beyond a stage at which the situation could have been corrected with limited impact.
  • an abnormal change in an animal's behaviour that the monitoring system did not capture, delaying veterinary examination.
  • a real rise in the temperature of a shipment that the system did not record because of a sensor fault or an interruption in data transmission, so the shipment advanced to a later stage of the chain without review.
  • a developing fault inside a pump or cooling unit that the predictive-maintenance system did not detect, so operation continued until a more serious stoppage occurred. The danger of this type lies in the fact that the user may not know that the system has erred. A false alert draws attention to itself, but the case the model missed may remain hidden until its effects appear. For this reason, error detectability is part of risk assessment, not merely the probability of its occurrence. In applications where an undetected case is highly dangerous, the system may be set to be more sensitive, that is, more inclined to refer suspected cases for inspection, even if that leads to an increase in false alerts. But this choice is not automatically correct; it must be verified that the human capacity for inspection can absorb the number of alerts, otherwise heightened sensitivity turns into congestion that strips the system of its usefulness.

Timing error: a correct answer that arrived after the decision window had passed The model may be correct about the type and magnitude of the problem, but its recommendation arrives at a time when action is no longer possible or worthwhile. This is timing error: the problem lies not only in whether the information is correct, but in where it falls within the agricultural timetable. Agricultural decisions operate within an intervention window: a limited period during which action can alter the outcome. After this window closes, even the best recommendations may become a description of what happened rather than a means of changing it. Examples include:

  • forecasting rainfall after irrigation has been completed, when the forecast can no longer save water or energy.
  • detecting an infection after it has passed the stage at which it could be contained by local inspection or limited intervention.
  • sending a frost alert after temperatures have already entered the critical range, even though protective measures require time to prepare and activate.
  • predicting fruit ripeness after the opportunity to arrange labour, packaging, or transport has been lost.
  • detecting a cooling fault after the shipment has left a point at which the fault could have been repaired or the route altered.
  • providing a price forecast after the sales contract has been concluded or the available options for storage and marketing have closed. For this reason, it is not enough for an evaluation study to ask, "Was the forecast correct?" It should also ask, "How far in advance of the event did it come? And was that lead time sufficient to carry out the action?" A slightly less accurate model may be more valuable if it provides an early warning that can be acted upon than a more accurate model that arrives after the opportunity to act has passed.

Magnitude error: the action is correct, but its quantity is inappropriate In some cases, the system recognizes the correct need but errs in the magnitude of the response. It may recognize that the crop needs irrigation, but estimate a quantity greater or less than required. It may determine that ventilation or cooling is necessary, but recommend a degree or duration that does not suit the actual conditions. This is magnitude error: the direction of the decision is correct, while its quantity, intensity, or duration is incorrect. The seriousness of this error appears in many applications:

  • Irrigation: too much may waste water and energy and increase leaching or root asphyxiation; too little may leave the stress unresolved despite irrigation having been applied.
  • Fertilization: the suggested nutrient may be appropriate, but the quantity or timing of application may not suit the soil or the crop stage.
  • Greenhouses: the need for cooling may be real, but an abrupt or prolonged change may create another form of stress or increase energy consumption.
  • Animal feeding: adjusting the ration may be justified, but the magnitude of the adjustment should not be considered apart from production status, health status, and the ration's full composition.
  • Storage and drying: lowering moisture or temperature may be required, but too much or too little operating time may affect quality and cost. The stakes are higher when chemicals, veterinary medicines, or withholding periods enter the decision. In these domains, the model's estimates alone must not be turned into a dose or an execution order. Reference must be made to the official label, local registration, the instructions of the competent authority, and a qualified specialist, with adherence to doses, pre-harvest intervals, re-entry periods, and other binding safety requirements.

Context error: information that is correct in one place and wrong in another Not every error results from a weak model. Sometimes the information is correct in the environment in which it was produced, but it is transferred to a user whose conditions differ in ways that make it inapplicable. The decision may change because of differences in:

  • the variety, breed, or life stage.
  • the soil type, its salinity, and its water-holding capacity.
  • the local climate, radiation intensity, humidity, and winds.
  • pest and disease pressure and their spread in the region.
  • the irrigation system, the available equipment, and the accuracy of the sensors.
  • the form and scale of the holding, and the availability of labour and expertise.
  • the prices of inputs and energy, and the cost of alternatives.
  • legally registered products, safety intervals, and regulatory restrictions.
  • the language in which the user understands the alert and the user's ability to act on it. An academic study may describe a crop response under specific experimental conditions, and the result may then be turned within an application into a general recommendation for all regions. An intelligent assistant may cite the name of a substance or practice from a foreign source, even though its registration or conditions of use differ in the user's country. A disease model may perform well on images taken in good lighting for a particular hybrid, then weaken when faced with another variety, a different phone camera, or symptoms overlapping with a nutrient deficiency. Scientific validity here does not mean refusing to transfer knowledge between environments; it means stating the limits of transferability and testing them. The question is not only, "Is the information correct?" but also, "Correct for whom, where, in which season, and under what conditions?".

Data and integration error: when the algorithm is innocent and the decision is wrong In digital agricultural systems, the model itself may be sound, but the data that reached it, or the way its result was conveyed, may be defective. This is an important category because it occurs at the interface between agriculture, software, and hardware. Examples include:

  • a sensor reading that was not updated, but the application displays it as though it were real-time.
  • an uncalibrated moisture sensor that sends consistent readings that are nevertheless incorrect.
  • a mix-up between units of area, volume, or mass when data move from one system to another.
  • the use of an incorrect time or time zone, so that the alert is associated with an unsuitable hour or day.
  • records are lost during a communications outage, and the calculation is then performed on incomplete data without alerting the user.
  • an operating command is repeated because of a software retry, even though the device has already executed the first command.
  • a field or crop record is linked to another record bearing a similar name.
  • the system continues to use an old model or recommendation after the variety, season, or production plan has changed.
  • a high confidence score is displayed in the user interface in a way that suggests certainty, even though the model has not been tested in this environment. These errors are not solved by retraining the model alone. They require unit verification, clear timestamps, sensor calibration, detection of missing data, prevention of duplicate commands, audit logs, and integration testing between software and hardware.

Individual error and large-scale error The consequence of an error is tied not only to its type, but also to the scale of its spread. A wrong recommendation seen by an agricultural engineer and then rejected is different from the same command when executed by an automated system across hundreds of valves or thousands of hectares. Automation gives error two contradictory powers:

  • it can execute the correct decision quickly, consistently, and at scale.
  • it can likewise repeat the same error quickly, consistently, and at that same scale. The wider the scope of execution and the less human intervention there is, the greater the need for operating limits, prior testing, gradual deployment, the possibility of manual shutdown, a log showing what happened, and a mechanism that returns the system to a safe state when data or communications are lost. The problem is not that the machine may make a mistake once; it is that a single software error may turn into a repeated practice before anyone notices.

The threshold is a professional decision, not merely a statistical number Models do not always say, of their own accord, "infected" or "healthy". Many of them produce a score expressing the strength of the signal or the model's confidence, and the system designers then set a threshold at which the output moves from silent monitoring to issuing an alert. If the threshold is lowered, the system becomes quicker to suspect:

  • It usually detects a greater number of true cases.
  • But it may also send a greater number of false alarms. If the threshold is raised, it becomes more conservative:
  • the number of false alarms decreases.
  • But it may miss true cases whose signal does not exceed the higher threshold. There is no ideal threshold suitable for all applications. The choice depends on what happens after the alert. If the application ranks images to determine which ones deserve to be examined first by the agricultural engineer, and if that examination is quick, inexpensive, and safe, then it may be acceptable to lower the threshold and refer a greater number of suspected cases. But if crossing the threshold leads directly to starting a machine, changing the environment of a greenhouse, excluding a shipment, or considering a chemical treatment, then the additional error may carry a cost that cannot be ignored. Nor should the "model confidence score" be understood as certainty. If the system displays a high score, this does not necessarily mean that the real-world probability of the decision being correct equals that score, unless the model has undergone appropriate calibration and its significance has been tested in the target environment. The interface should therefore explain the limits of the score, not adorn the recommendation with a number suggesting a degree of certainty it does not possess.

How does the threshold change with the decision?

Type of useWhat happens after the alert?Reasonable trade-offNecessary safeguard
Initial image-screening toolThe image is referred for human reviewAdditional alerts may be accepted in order to reduce missed serious casesA statement that the result is an initial suspicion, not a final diagnosis
Irrigation monitoringThe user reviews the reading, the weather, and the condition of the soilBalancing missed stress against the cost of over-irrigationDisplaying the data source and its recency, and allowing manual verification or override
Disease or pest alertThe specialist visits the site or requests an additional image or sampleRaising sensitivity in rapidly spreading cases while managing the number of alertsA clear path for human or laboratory confirmation
Greenhouse controlThe operation of ventilation, cooling, or heating may changeGreater strictness because the decision directly affects the environment, production, and energySafe operating limits, gradual changes, and the possibility of manual stop or override
Robot or field machineDetection turns into movement or physical actionVision accuracy alone is not enough; safety and the full end-to-end action must be testedSafe stop, obstacle detection, and limits on speed, force, and operating zone
Sensitive chemical or veterinary decisionIt may result in a treatment, dose, or withholding periodA model alert must not become a standalone execution orderSpecialist review and reference to the registration, the official label, and the legal requirements

Risk is not measured by the probability of error alone An error may be rare yet catastrophic, or frequent yet limited in effect. Risk assessment therefore requires looking at several dimensions together:

  1. Probability of occurrence: How often is this error expected to occur?
  2. Severity of consequence: Does it cause limited inconvenience, financial loss, environmental harm, or a safety hazard?
  3. Scale of exposure: Does it affect a single plant, an entire field, a herd, or a supply chain?
  4. Speed of harm progression: Are there hours or days available for intervention, or does the error quickly become irreparable?
  5. Detectability of the error: Does the user see it immediately, or does it remain hidden until the loss appears?
  6. Reversibility: Can the decision be easily undone, or can its effect not be reversed after execution?
  7. Degree of automation: Does the system propose and leave the decision to the human, or does it execute the action directly?
  8. Affected parties: Who reaps the benefit, and who bears the harm if the system is wrong? For that reason, a wrong alert on a screen may be less dangerous than an apparently correct command executed automatically with the wrong unit of measurement. And a small repeated error in estimating irrigation may be more consequential over a whole season than a large error that occurred once and was quickly detected.

Before adopting the threshold or permitting execution The programmer should not choose the threshold alone; they know the behaviour of the model, but may not know the full consequences of the agricultural decision. Nor should the agricultural specialist determine it without data on the types of errors and the system's ability to express uncertainty. It is a meeting point between agricultural knowledge, software engineering, economics, safety, and regulation. For this reason, the following questions should be answered clearly:

  • What is the condition that the system is trying to detect or predict?
  • What data does it rely on, and how do we know that they are current and correct?
  • Which is more harmful in this use: the false alarm or the case that went undetected?
  • Who receives the alert, and do they understand its meaning and its limits?
  • What action is expected after the alert?
  • Is the action merely an inspection, or a physical, chemical, or financial intervention?
  • What is the cost of carrying out the action if the alert is wrong?
  • What is the harm if the system does not issue an alert even though the problem is present?
  • Does the alert arrive before the time available for intervention runs out?
  • Can the action be reversed after it has been carried out?
  • How many cases can the human team review without exhaustion?
  • What does the system do when the data are incomplete or contradictory?
  • Can it refrain from making a recommendation and declare that it does not have sufficient information?
  • Is there a safe manual alternative when the connection, sensor, or service fails?
  • Does the system keep a record showing the data, the recommendation, the action, and who approved it?
  • Does the decision require review by a specialist or by a legal or professional authority?
  • Was the full system tested, from data collection to execution of the action, or was the model alone tested?

Responsible design does not promise to eliminate error; it prevents error from turning into harm There is no model free of error, just as no human decision is free of it. The realistic aim is not to claim the elimination of uncertainty, but to design a system that knows where it may err, detects error early, limits the spread of its effects, and gives the human operator a clear opportunity to intervene. This is achieved, according to the level of risk, through means such as:

  • Displaying the degree of confidence and its limits in clear language.
  • Requesting an additional image or reading when the data are insufficient.
  • Refraining from issuing a definitive recommendation in unfamiliar cases.
  • Separating the initial alert from the diagnosis or final decision.
  • Referring sensitive cases to a specialist.
  • Using more than one data source instead of relying on a single sensor.
  • Setting safe limits on commands that can be executed automatically.
  • Requiring human approval before high-impact decisions.
  • Providing a manual stop and a safe operating mode in the event of failure.
  • Recording data, recommendations, and commands so that investigation and review are possible.
  • Monitoring changes in performance across seasons, locations, and varieties.
  • Testing interruptions, stale data, unit conflicts, and duplicate commands, rather than testing the model only under ideal conditions. The model does not itself decide which errors can be tolerated; that is an agricultural, economic, ethical, and safety-related trade-off in which human expertise and technical knowledge both have a role. The true measure of a system's maturity is not merely that it errs infrequently, but that it knows how to behave when in doubt, how to stop when the data are insufficient, and how to prevent a limited error from becoming widespread harm that cannot be reversed.

8. How Do We Compare Two Applications, Not Two Numbers?

Comparing the higher number with the lower is not enough, because applications may differ in task, data, context, maturity, and mode of operation. A rigorous comparison begins by standardizing the question before comparing the answer. An evaluation sheet for an agricultural application

DimensionQuestion to be answered
Identity and timeWhat is the name of the system, its version, and the date of its examination?
TaskWhat agricultural decision is it trying to improve?
UserWho uses it, and with what skill and authority?
MaturityIdea, prototype, validation, field trial, operation, or measured impact?
DataWhere did they come from, and do they represent the target crop, location, and user?
ComparisonWhat baseline or alternative was it compared against?
MetricHow was it defined, what is its unit, and what does it not measure?
ValidationInternal, external, local, or field-based?
ErrorsWhat types of error are there, and who bears their consequences?
OperationWhat happens in the event of interruption, corrupted data, or changing conditions?
TimingDoes the recommendation arrive before the window for intervention closes?
ImpactDid an agricultural, economic, or safety outcome change?
CostWhat is the total cost of ownership, operation, and training?
RightsWho owns the data, and how are they exported or deleted?
ResponsibilityWho approves the decision, who stops it, and who remedies the harm?
Conflict of interestWho funded, described, and evaluated the system?

A field for which information is unavailable is not filled with guesswork or with an average taken from another product; it is written clearly: "Unknown". Missing information is part of the evaluation result, not a formal defect to be concealed. An illustrative example: 97% or 89%? Let us suppose there are two applications for classifying images of plant diseases:

  • The first application claims an accuracy of 97% within an image database selected from a single source.
  • The second application claims 89% in a test that included different farms and devices, offers an "unknown" category, abstains when image quality is poor, and refers sensitive cases to a specialist.

It is not permissible to declare the second application superior merely because its description appears more responsible, just as it is not permissible to choose the first because its number is higher. Comparison requires knowledge of the task, the distribution, the type of error, the baseline, the cost, and the operating conditions. But the example reveals a basic rule: the higher number does not compensate for the absence of evidence on transferability, safety, and operation. An application with the lower overall metric may be more suitable for a particular task if the pattern of its errors is known, its limits are declared, and its escalation path is operationally feasible. Sound comparison does not ask which of the two numbers is larger. It asks which system, in this context, supports the most useful and least risky decision with the most appropriate evidence.

9. What Does Not Appear in the Demonstration and on the Product Page?

The demonstration is important; it shows that the system was able to perform a task under certain conditions. But the demonstration does not usually choose the worst day of the season, the weakest connection, the most ambiguous image, or the least experienced user. Crucial information may be absent from the demonstration or the summary, including:

  • Cases excluded before testing.
  • Poor-quality images or rare classes.
  • The number of connection interruptions.
  • The time required for data cleaning and preparation.
  • The extent of human intervention involved in producing the result.
  • The number of false alarms.
  • Cases in which the user refrained from implementing the recommendation.
  • The cost of calibration and maintenance.
  • Failures that appeared after the trial ended.
  • The difference between the best site and the average across all sites.
  • The effect of changes in season, device, or data source.
  • The number of users who stopped using the system.
  • The cost of taking the system from a prototype to a stable service.

One may say that the system "reduced water consumption", without stating:

  • Compared with which practice?
  • Did water use decrease per hectare or per tonne?
  • Did total water withdrawal from the source decrease?
  • Were productivity and quality maintained?
  • Was the season rainier?
  • Did the cultivated area expand?
  • What happened to energy consumption?
  • What is the cost of the system compared with the value of the savings?

One may also say that an application is "AI-powered", when in fact the algorithm is only a small part of a service whose success depends on the quality of measurement, connectivity, support, maintenance, and the user's response. Signs that warrant pause Judgment should become more cautious when one or more of the following patterns appear:

  • A percentage with no denominator or definition.
  • Accuracy with no description of the test set.
  • A comparison with a weak system or with no intervention at all.
  • Images of success with no presentation of errors.
  • A short trial presented as though it were sustained operation.
  • A model result presented as though it were an economic impact.
  • A vendor description presented as if it were an independent evaluation.
  • The word "saving" without reporting total consumption.
  • The word "instant" without a measured response time.
  • "Multilingual" without testing user comprehension.
  • "Interpretable" without stating what is explained and to whom.
  • "Safe" without defining the safe state or the failure-response plan.

A demonstration proves that the system was able to operate once under selected conditions. An agricultural service, by contrast, proves that it can operate when it does not choose the conditions itself.

10. The Evidence Required Is Proportional to the Risk of the Decision

Not every digital function requires the same degree of proof. Reminding a user of a crop inspection date does not carry the consequences of a chemical recommendation or a credit decision. But the low cost of the application does not reduce the consequences of an error if the decision is sensitive.

Degree of riskExamplesMinimum required
LowSorting images for inspection, a reminder to visit the field, organizing a recordFunctional verification, ease of correction, and clear indication that the output is only an aid
MediumYield prediction, preliminary irrigation scheduling, crop sorting, maintenance alertLocal verification, baseline, operational monitoring, manual alternative
HighDiagnosis leading to a sensitive intervention, automatic control, a decision affecting food safetySpecialist verification, human approval, a decision log, defined limits, abstention, and a safe mode
Regulated or highly consequentialDosages, pesticides, PHI, REI, legal registration, veterinary decisions, credit or insuranceA valid current official source, specialist review, the right to object, clear responsibility, and prevention of unauthorized execution

Note: Here, baseline means the reference against which the performance of the artificial intelligence system is compared to determine whether it delivered a real benefit. Saying that the system "performed well" is not enough unless we know: well compared with what? The baseline may be one or more of the following:

  • Current practice before using the system: such as the farmer's usual method for estimating irrigation timing or sorting the crop.
  • Results from a previous season or period: such as the average field yield across previous seasons, while taking differing conditions into account.
  • A simple non-AI method: such as a fixed irrigation schedule, a clear agronomic rule, or a historical average.
  • An estimate by a human expert: such as the agricultural engineer's yield forecast or the result of manual sorting.
  • An existing conventional model: if the project is replacing a system or model already in use.

Example: yield prediction If an AI model predicts crop yield, it is not enough to say that its prediction was close to the actual outcome. It should be compared, for example, with the field's average yield in previous years or with the estimation method currently used by the agricultural engineer. If the model is not better than this simple baseline, or does not provide its prediction at a more suitable time, its usefulness may not justify the cost of data collection and system operation. Example: irrigation scheduling If the system recommends the timing and quantity of irrigation, the baseline is the usual irrigation schedule, or the farmer's decision based on inspection and experience, or an approved guidance rule. The outcomes are then compared: did the system reduce total water and energy consumption? Did it preserve yield and quality? Did it reduce water stress? And did the benefit remain after accounting for the cost of sensors, connectivity, and maintenance? Example: crop sorting The baseline may be the current manual sorting process. The comparison is not limited to the system's speed; it also includes the proportion of products classified incorrectly, consistency of quality, the extent of damage, operating cost, and the worker's ability to review ambiguous cases. Example: maintenance If the system issues early alerts about faults in pumps or cooling units, the baseline may be the routine maintenance schedule or the number of faults and stoppages before the system was installed. The aim is to determine whether the alerts actually reduced downtime and losses, not merely to know how many alerts the application generated. So, the baseline is not a minimum threshold for acceptance, nor a programming starting point; rather, it is a documented picture of what was happening before the system, or of what a simpler alternative can achieve. Without it, we may attribute to AI an improvement caused by the weather, changing prices, worker experience, or differences between seasons. If the technical term is retained in the table, the clearest wording is: Local verification, comparison against a clear baseline — that is, against current practice or a simple alternative — continuous monitoring during operation, and provision of a safe manual fallback. Low-risk functions A greater degree of flexibility may be accepted if the error is easy to detect and reverse, and the tool does not execute a harmful action. But even here, the platform should not conceal its limits or collect data it does not need. Medium-risk functions These require comparison against a realistic baseline, testing on local data and under local conditions, and an understanding of what happens if the system fails or the connection is interrupted. The recommendation may not be dangerous in itself, but it may become costly if used over a wide area. High-risk functions These should not move directly from model prediction to execution. They require:

  • identifying the person authorized to approve.
  • showing the source and the limits.
  • preventing an answer when data are insufficient.
  • recording the recommendation and the human intervention.
  • providing an escalation path.
  • testing the safe mode.
  • separating general information from professional or legal judgment.

Legally regulated functions The algorithm must not replace the official label, the competent authority, the veterinarian, or the authorized engineer. It may help collect information, retrieve documents, and flag inconsistencies, but it cannot confer legal registration or professional authority that it does not possess. The governing rule is that the degree of verification and oversight should increase with the consequence of error and the difficulty of reversing it.

11. Three Case Studies in Reading the Evidence

The following cases are analytical teaching examples, not the results of new experiments. Their purpose is to show how to apply the chapter's method. Case One: A Phone Application for Classifying a Plant Disease The application displays a picture of the leaf and returns the name of a possible disease and a confidence score. A quick reading would say: the application diagnoses the disease. A disciplined reading, however, breaks the claim down:

  1. Does the model recognize a visual pattern within specified classes, or does it provide a full causal diagnosis?
  2. Is the image dataset representative of the target varieties and environment?
  3. Was it tested on phone images in the field?
  4. Can it detect that the case falls outside the training classes?
  5. Does it ask for information about the plant's age, location, and the spread of symptoms?
  6. What does it do when confidence is low?
  7. Does it display official sources?
  8. Does it propose a treatment, and who approves it?
  9. Does it distinguish between the disease name, its severity, and the stage of its spread?
  10. Can the user correct the result and send it for review?

The application may be very useful in triaging images and identifying the cases that merit inspection, without being a substitute for professional diagnosis. Describing the function accurately does not diminish its value; rather, it prevents its use in a context for which it was not designed. Case Two: A System That Claims a 30% Saving in Irrigation Water Let us assume, for the purpose of analysis, that a platform claims a 30% reduction in irrigation water. Before accepting the statement, we must ask:

  • Is the percentage per hectare, per tonne, or for the farm as a whole?
  • What is the baseline: a fixed schedule, the farmer's practice, or another system?
  • Did the area remain constant?
  • Was the season comparable in rainfall and temperature?
  • Did productivity and quality remain unchanged?
  • Was water consumption shifted to another stage?
  • What happened to the pumping energy?
  • Did the result include outage days?
  • How much did the sensors, connectivity, and maintenance cost?
  • Did the saving persist for more than one season?

If water use per tonne decreases as the cultivated area expands, total abstraction may rise. If water use declined because the season was rainier, the entire result is not attributable to the system. And if the saving was achieved with a large decline in quality, the result may be economically unacceptable. The percentage should not be dismissed, but it is incomplete without its denominator, its baseline, and the outcomes that accompany it. Case Three: A Generative Agricultural Advisory Assistant The assistant answers, in natural language, a question relating to disease, irrigation, or fertilization. Its success is not measured by the elegance of the sentence alone. The following should be examined:

  • Did it understand the crop, the region, and the stage?
  • Did it retrieve a recent reference from a trusted source?
  • Did it convey the source's conditions accurately?
  • Did it confuse two countries or two legal registrations?
  • Did it distinguish between probability and diagnosis?
  • Did it refrain from inventing the dose or the pre-harvest interval?
  • Did it ask for the missing information?
  • Did it escalate the case to a specialist when necessary?
  • Did the farm's data remain protected?
  • Can the user see the source, object to the answer, and correct it?

The assistant may be excellent at explaining probabilities, gathering questions, and organizing evidence, yet dangerous when it provides a sensitive figure with confidence it cannot justify. Maturity here does not mean answering everything; it means knowing when an answer is useful and when referral is the responsible answer.

12. The Reader's Protocol Before Trusting a Digital Agricultural Claim

The following questions may be used when reading a study, a product presentation, or a chapter of this book:

  1. What is the specific agricultural task?
  2. Is it detection, classification, prediction, recommendation, control, or execution?

  3. What type of claim is it?
  4. Does it concern the existence of the product, a feature, technical performance, or agricultural and economic impact?

  5. Who issued the claim, and who evaluated it?
  6. Is the speaker the developer, an independent body, researchers, or users?

  7. What is the source of the data?
  8. And what are the crop, the region, the season, the device, and the sample unit?

  9. What is the baseline?
  10. Was the system compared with current practice, an expert, a simple rule, or no intervention?

  11. What is the metric?
  12. What does it measure, what does it not measure, and is it suited to the consequence of error?

  13. Was there independent or local validation?
  14. Or is the result still confined to the developer's data or the original study?

  15. Which type of error is most dangerous?
  16. And who bears its cost or harm?

  17. Does the result arrive in time?
  18. And can the user actually act on it?

  19. What happens when failure occurs?
  20. What happens if the network goes down, the sensor fails, or an unknown case appears?

  21. Did the final outcome improve?
  22. After accounting for yield, quality, water, energy, labour, cost, and risk?

  23. What requires local validation before reliance?
  24. And what limits of use must be declared?

After answering, the claim may be placed in one of four states:

  • Supported within a known context: Fit for use within the declared limits.
  • Promising and in need of local validation: There is good evidence, but transferability has not been demonstrated.
  • Described by its provider: We know what the provider claims, but we do not yet have sufficient independent evaluation.
  • Insufficient to support the decision: The evidence is incomplete, unsuitable, or disproportionate to the risk of use.

The fourth state does not mean that the technology is worthless; rather, it means that the gap between potential and justified reliance has not yet been closed.

Chapter Summary

Knowledge does not become trustworthy because its source carries a famous name, nor does an application become mature because its model achieved a high score. Trust is built when we know what was measured, on what data, against what comparator, and under what conditions; then trace the path from model performance to system operation, from recommendation to action, and from action to its effects in the field, the economy, and human wellbeing.

A figure may be correct within its study, the product may exist in the marketplace, and the algorithm may be advanced in its design, yet none of these is sufficient evidence that the decision has become better on a particular farm. Here the method serves its purpose: not to dim the promise of technology, but to turn it from quick fascination into testable knowledge, and from knowledge into a defensible decision. Scientific rigour does not ask every application to prove everything from its first day. It asks it to declare what it has in fact proved, to distinguish between the level it has reached and the level to which it aspires, and to increase the degree of validation whenever the consequence of error rises. When the reader learns to ask about the task, the metric, the context, the error, the maturity, the transferability, and the impact, the reader is no longer captive to the highest number or the glossiest presentation. The reader becomes able to distinguish between a model that succeeded in testing, a system that was able to operate, an application that changed the decision, and a technology that proved it deserves its place in agriculture.

Evidence Notes

Sections One and Two describe how the book was constructed and how its claims were disciplined. [SRC004] was used as an example of a published systematic review, not to confer its methodological status on this work. [SRC006] was used to clarify that the coefficient of determination is limited to the study set-up and must not be turned into a general "accuracy percentage", and [SRC012] as an example of the need to confine computational results to the experiments from which they came and not turn them into a commercial or field guarantee. The discussion of risk, oversight, and the operational cycle is guided by the AI Risk Management Framework [SRC033]. The maturity ladder, assessment cards, risk levels, and case studies are an editorial teaching synthesis intended to help the reader evaluate the evidence; they are neither an official regulatory standard nor the results of new experiments conducted for this book.

Interactive learning lab

Build an evidence chain for agricultural AI

Follow the path from observation to a claim that can be checked, reproduced, and challenged.

This enrichment complements the chapter and does not replace its editorial text.

Evidence chain

Confidence grows when every transition remains visible and reviewable.

  1. Observe with context
  2. Preserve provenance
  3. Validate and compare
  4. State a bounded claim

Put this chapter into practice

Choose a situation to see what evidence to check and the responsible next step.

Choose a situation to see what evidence to check and the responsible next step.

Evidence maturity levels

Showing 3 of 3 rows.
Evidence maturity levels
LevelWhat it containsWhat it can support
Raw observationOriginal value, time, place, sourceTraceable inspection
Validated evidenceChecks, units, context, uncertaintyReasoned comparison
Decision-grade claimEvidence, limits, responsible reviewerA bounded recommendation

Check your evidence judgement

Choose an answer to receive immediate feedback.

Question 1 What does provenance answer?
Question 2 Does a high-confidence model output automatically become decision-grade evidence?
Question 3 Why should a claim be bounded?
Score: 0 of 3 correct.

Reader community

Comments and scientific reviews

Contributions are linked to this language and section. Nothing appears publicly until an authorized editor approves it.

Approved contributions

No approved contributions have been published for this chapter yet.

Submit a contribution

Every submission is checked for relevance, safety, and scientific clarity before publication.

Your name, email address, contribution, and book-section context are stored on this site for moderation. Your email address is not displayed publicly, and this book does not retain your IP address or browser identifier with the contribution. Do not include passwords, API keys, phone numbers, or other sensitive personal data.

Only aggregate events are counted. Search terms, comment text, private notes, email addresses, IP addresses, and user-agent strings are never stored in book analytics.