Skip to book content
Menu
International Office · Istanbul, Türkiye dr.alaa@aladdin.my.id +90 541 514 37 21

Aladdin Interactive Book

Chapter Four: Artificial Intelligence Methods in the Agricultural Context

0%
Book knowledge toolsSearch and discuss Part ISearch this language edition or ask Chat V2 a question grounded in the approved book passages.
Discuss with Chat V2Book-only mode is the default. Every supported answer must cite an actual indexed passage.From this book only

Part Two — Foundations of Artificial Intelligence and Data

The question of this part: what happens inside an artificial intelligence system from the moment it receives data until its answer or recommendation reaches the farmer?

When an intelligent system proposes a time to irrigate, identifies a disease affecting a plant leaf, or predicts production, the answer may seem to have appeared all at once on the screen. Yet that answer is the final stop on a long journey the user does not see: it begins with data collected from the field, satellites, or past records, then passes through measurement, cleaning, classification, and analysis before becoming a prediction, recommendation, or warning. In public discussion, artificial intelligence is sometimes presented as though it were a single mind capable of understanding and making decisions in every situation. In reality, it is not one tool, but several families of methods and models, each designed for particular tasks. The tool that analyses an image of a plant leaf is not the same as the tool that predicts crop yield; and a system that extracts information from thousands of documents does not work in the same way as one that determines when to irrigate. Every method has strengths, assumptions on which it depends, data it requires, and potential errors that must be understood before its results are trusted. This part is not intended to turn the reader into an artificial intelligence specialist, nor to burden them with mathematical and programming detail. Its purpose is deeper and more practical: to give readers an understanding that protects them from being dazzled by a result before knowing how it was produced, and helps them distinguish between a system built on appropriate data and clear evidence and one that offers confident answers unsupported by sufficient knowledge. To that end, we will try to look inside the system, in language that combines precision with simplicity. We will ask: what data did it learn from? Where did those data come from, who collected them, and in what land, climate, and season? Do they represent the conditions of the farm where the system will be used, or a different environment? What can the model see, and what remains outside its view? How does it know that its answer is correct? And how might it err, hesitate, or express more confidence than the evidence allows? These are not technical questions remote from agricultural reality. If weather data are incomplete, a satellite image is obscured by clouds, a sensor is uncalibrated, or the model was trained on a different crop or climate, the defect may pass from one stage to the next until it reaches the farmer as an apparently precise recommendation. This does not mean that artificial intelligence is useless; it means that its true value comes not from its computational intelligence alone, but from the quality of the entire chain that produced its answer. We begin with artificial intelligence methods and the logic of choosing each one: why is one tool suited to a problem while another is not? We then move to data, because they are the material from which the system learns and the mirror through which it sees the world. After that, we examine data provenance and quality, so that we know where the information originated, the path it took, and what changes were made to it. We then reach remote sensing, where images from satellites and drones, together with field devices, are turned into indicators of plants, soil, and water. From there we move to decision-support systems and knowledge retrieval, to see how measurements, models, and agricultural expertise are brought together in an answer that can be used at the right time and place. This order is not arbitrary; it reflects the path knowledge takes through the system. An unsuitable method may corrupt the inference, weak data may undermine the best models, an unknown source makes the result difficult to verify, and an inaccurate measurement produces an inaccurate recommendation. An answer that does not make its basis and limits clear may turn from a tool of assistance into a new source of risk. Understanding these foundations does not mean that the farmer must know every equation or line of code, any more than a tractor driver needs to be a mechanical engineer. But the farmer needs to know enough to distinguish sound operation from malfunction, real capability from a marketing promise, and a tool that helps with decision-making from one that asks the farmer to surrender the entire decision to it. In this part, we do not view artificial intelligence as a mysterious box that issues answers, but as an interconnected chain of data, assumptions, measurements, models, and human decisions. Every link in that chain affects the one that follows. The most important question, therefore, is not only: What did the system say? It is also: How did it arrive at what it said, on what evidence, and within what limits can it be trusted? The discussion remains concrete, without reducing its scientific depth or turning the introduction into a heavy technical explanation.

Chapter Four: Artificial Intelligence Methods in the Agricultural Context

Chapter Message

There is no single method called 'agricultural AI' that suits every crop, farm, and decision. What we call artificial intelligence is a broad range of methods: some learn from past examples, some search for unusual patterns, some interpret images, some track change over time, some compare alternatives under constraints, and some retrieve knowledge and formulate it into an answer. The value of any method does not begin with the fame of its name or the novelty of its model, but with its fit to the agricultural question we want to answer. A system error may be nothing more than an image that needs to be retaken; or it may be a delayed decision that misses a treatment window, or an automated operation that puts a crop or a worker at risk. Selection therefore begins with the decision and the consequence of error, and only then returns to the data and the method—not the other way round. This chapter does not ask the farmer to become a programmer, nor is it content merely to introduce the informed reader to familiar names. Its aim is to build a bridge between what happens inside the model and what happens afterwards in the field: what does the system see? What does it learn? What does it output? Where can it go wrong? And when is the result supportive information, and when does it become an action that requires oversight and responsibility? Before the names of algorithms: the journey of the question through the system Let us imagine a farmer who finds spots on tomato leaves, takes a picture, and asks an app: 'What is the problem?' The task seems simple, but the system does not see the leaf as a person sees it, nor does it know on its own the field's history, the variety, the humidity, or how widely the symptoms have spread. It receives a digital representation of the image, compares it with patterns it has learned, and then produces a class, a probability, or a set of probabilities. After that, another application may turn the result into a recommendation. Between the image and the recommendation lies an entire chain of choices. Any intelligent agricultural system can be understood through six stages:

  1. The question: What problem or decision do we want to support?
  2. The inputs: What data can the system actually see?
  3. The representation: How has agricultural reality been turned into an image, a number, a text, or a time series?
  4. The method: What computational family links the inputs to the output?
  5. The output: Is the result a classification, an estimate, a probability, a ranking, or a recommendation?
  6. The action: Who will use the result, and what happens if it is wrong or delayed?

Separating these links matters. A model may be good at classifying an image, while the image itself is not representative of the field. A prediction may be reasonable, but turning it into a dose or a timing requires agricultural rules and safety constraints that the model has not learned. In mature systems, these distinctions do not disappear behind the word 'intelligent'; they are shown to the user, so that they know where trust is warranted and where verification is needed.

1. From the Question to the Representation

The agricultural question as a person frames it is usually broader than the question a model can handle. Asking, 'Is this plant all right?' combines in a single phrase the possibility of disease, nutrient deficiency, water stress, mechanical damage, and the effects of weather. A model, by contrast, needs a specific task whose inputs can be represented, whose output can be defined, and whose error can be measured. Take the example of an affected leaf. It can be turned into different problems:

What do we ask of the system?How does the result appear to the user?What is its practical value?
Determining the image's general conditionA message such as: 'The leaf appears healthy', or 'There are signs of concern', or 'The image is insufficient for judgement'Helps with initial triage and indicates whether better imaging or specialist review is needed
Locating the suspected signThe system places a visible circle or box around the spot or part of the leaf that drew its attentionHelps the user confirm that the system was looking at the spot itself, not the soil, a shadow, or the image background
Delineating the boundaries of the affected areaThe system colours the damaged area within the leaf or draws its boundaries clearlyHelps estimate how much of the leaf is affected and compare it with later images
Comparing images taken at different timesThe system shows whether the affected area has expanded, receded, or remained the sameHelps monitor how the condition is developing instead of relying on a single image
Ranking possible causesIt presents several possible causes and explains what supports each one and what information is missingDirects the user to an additional image, weather information, or a field inspection before a decision is made

These are not multiple names for the same task. Each question requires data suited to it, correct answers from which the system can learn, a way to measure its errors, and a clear step that follows once the result appears. Classification may be sufficient for initial triage, but it is not enough to guide a robotic tool to a precise location. And a damage map may be visually excellent, but it does not prove the cause of the damage. The same applies to weeds. If all that is required is to know whether a weed is present, a classification task is enough. If what is required is to calculate weed density, we need counting or area estimation. If what is required is to guide a mechanical tool, we need a precise location, response time, a control system, and a safety barrier. The shift from 'it saw the weed' to 'it moved towards it' is a shift from information to action, not a minor programming detail. That is why the unit of decision comes before model selection. Does the decision concern a single leaf, a plant, a row, a sector, or an entire field? And are we predicting a field's yield before harvest, or the total output of a region for planning storage and transport? If the unit changes, the meaning of the data and of the error changes with it. A good estimate at regional level may not be suitable for deciding irrigation in a small sector, even if both use the word 'yield' or 'production'. As for representation, it is the simplified digital picture of reality that the system allows. It may represent the soil by a moisture measurement at a single depth, or represent the plant by an overhead image, or represent field management by a record of dates. Every representation illuminates one aspect and leaves other aspects outside view. A model does not know what has not been fed into it, and it cannot correct a vague definition of the objective merely by increasing its computational power. A practical question: Before asking about the name of the algorithm, ask: what is the unit of decision? What does the system see? What does it not see? And what action will be built on its output?

2. Supervised Learning: Learning from Examples with Known Outcomes

Supervised learning is an approach in which the system learns from examples with known outcomes. If we want to train it to distinguish between healthy and affected plant leaves, we show it images that a specialist has already reviewed and assigned a condition to. If we want to train it to predict yield, we provide data on weather, soil, variety, and field management, and with each record the amount of yield actually measured after harvest. From a large number of such examples, the system tries to learn the relationship between the data and the known outcomes, and then uses what it has learned to estimate the outcome for a new case it has not seen before. The known outcome is called a expert-verified reference class when it is the name of a class, such as 'healthy' or 'affected', and a reference value when it is a measured number, such as the amount of yield. A model does not usually begin with a list of rules written by an expert, such as: 'If a spot of this colour appears, then the condition is such and such.' Instead, the developer presents it with a large number of examples with known outcomes, and it searches on its own for the features that recur with each outcome. When it sees a new example, it compares it with the patterns it has learned and then estimates the closest outcome. This capability is useful, but it carries an important risk: the system may learn an easy side cue instead of learning the agricultural feature that matters to us. It does not know on its own which part of the image represents the disease and which part is merely background; rather, it learns from the way the examples were collected and presented to it. Let us imagine that we want to train a system to distinguish between a healthy plant and an affected one. All the healthy plants were photographed inside a greenhouse with a white wall, while all the affected plants were photographed in an open field with dark soil visible behind them. What we want is for the system to learn the shape of the symptoms on the leaves, but it may discover an easier route: the white background accompanies the healthy images, and the dark soil accompanies the affected ones. The system may seem successful when tested on images collected in the same way, because the backgrounds in the test resemble the backgrounds it learned from. But it may fail on a new farm—for example, by classifying an affected plant inside a white greenhouse as healthy, or by flagging a healthy plant because dark soil lies behind it. The system did not set out to mislead; it learned the easier association the data offered it, rather than the disease signs the designer intended. The same problem arises in numerical data. If a harvest sensor gives readings above or below reality because it is uncalibrated, the model will learn from incorrect numbers. If the output of some fields is recorded in tonnes per hectare, while that of other fields is recorded as total tonnes without clarifying the difference, the system will treat incomparable values as if they were one and the same. And if a single name is used for different disease conditions, the model will not know the correct boundaries between them. In each of these examples, the error does not remain outside the model; it enters the very material from which the model learns. The algorithm may perform its computational task accurately, yet still learn a relationship built on a faulty measurement or an unstable definition. An advanced method cannot extract from data a certainty that was not present there to begin with. So it is not enough to find a written answer next to every image or record. We need to know:

  • Who determined the reference condition or reference value, and what rules did they follow?
  • Was the result confirmed by a single specialist, reviewed by more than one specialist, or established by a reference test?
  • Was the confidence level recorded, and was an 'inconclusive' result allowed when the evidence was insufficient?
  • Did the definition of each condition remain consistent across farms, institutions, and seasons?
  • Do the examples include easy and difficult conditions, early and advanced symptoms, and mixed cases?
  • Do the images and records represent diversity in varieties, locations, devices, and lighting?
  • Were closely similar images or records separated between training and testing, so that the system would not be tested on highly similar versions of what it had seen before?

Allowing the system to say, 'There is not enough information to judge,' is not a sign of weakness. In agricultural applications, disciplined abstention may be safer than forcing the system to choose a disease name for every image. A responsible result does not merely state the most likely possibility; it also makes clear the limits of confidence, the missing information, and the next step required for verification. For the advanced reader: a large number of examples does not compensate for narrow diversity or weak reference results. A small independent set representing different farms, seasons, and devices may be better able to reveal the model's limits than thousands of similar images. Therefore, the diversity of environments, the consistency of the outcome definition, the independence of the test data, and the cost of each type of error must be evaluated alongside the size of the dataset.

3. Unsupervised and Semi-Supervised Learning: What Do We Do When Documented Outcomes Are Scarce?

In the ideal case, we may wish for a specialist to review every image, reading, or record, and to determine the correct outcome associated with it. But this is rarely possible in reality. A farm may collect thousands of images and millions of sensor readings, while the time, money, or expert capacity needed to review every record individually is not available. When documented outcomes are few or unavailable, methods can be used to help organise the data and discover recurring patterns within it, without starting from a correct answer attached to every example. Among the most important of these methods are unsupervised learning, anomaly detection, and semi-supervised learning. These methods are similar in that they make use of unreviewed data, but they do not perform the same function and do not confer the same degree of certainty on their results. First: unsupervised learning In unsupervised learning, the system is not given a prior list telling it, 'This zone is good,' 'this one is weak,' and 'this one has a drainage problem.' Rather, it searches within the data for records that resemble one another in their characteristics, and then groups the similar ones together. Let us suppose that we have readings from different zones in the field, including soil moisture, productivity, temperature, and reflectance characteristics captured by aerial images. The system may find that some zones resemble one another in these readings and place them in one group, while placing other zones in different groups. This result may help the farmer to see a difference that was not previously clear, but it does not automatically explain its cause. Zones that share low yield may be mathematically similar, but the reason for the decline may differ from one place to another:

  • One zone may suffer from poor drainage.
  • Another may be affected by shade.
  • A third may differ in soil type.
  • And a fourth zone's reading may be low because of an uncalibrated sensor.

So, at this stage the system says: 'These records are similar according to the data I analysed.' It does not say: 'I have proved that this zone is poor or that it needs a specific quantity of fertiliser.' This difference is fundamental. Computational grouping helps direct attention, inspection, and sampling, but it does not replace agricultural interpretation. Nor do the groups found by the system exist independently of the designer's choices. The result is affected by the variables entered into the analysis, by the way their units are standardised, by the number of groups to be formed, and by the degree of similarity that the algorithm considers important. If we omit drainage data, for example, the system will not be able to use drainage in explaining the difference, however advanced the method of analysis may be. Second: detecting unusual cases One important use of these methods is detecting unusual cases. The system learns the usual pattern of readings or behaviour over a certain period, then alerts the user when a case appears that differs from it to a substantial degree. The difference may be:

  • A moisture sensor whose readings have begun to drift gradually.
  • A sudden increase in water consumption that may indicate a leak or a change in operation.
  • An unusual drop in the flow of an irrigation line.
  • A noticeable change in an animal's movement or feed-intake pattern.
  • A temperature reading that remained constant for an unusually long time, which may indicate sensor failure.

But the alert alone does not prove that a fault or disease has occurred. It means only: 'This case differs from what the system regarded as usual, and therefore warrants examination.' The cause may be a real problem, a natural seasonal change, a new agricultural operation, or a measurement error. Therefore, a signal of difference must not be turned directly into a diagnosis or an operational instruction. Nor is a large number of alerts evidence of the system's quality. If the system sends, for example, one hundred alerts a day to a team that cannot examine more than five, the tool will become a source of noise, and important warnings may be lost amid repeated alerts. Therefore, the system should be calibrated according to three factors:

  • The seriousness of the condition it may fail to detect.
  • The number of false alarms that the team can handle.
  • The ability of the worker or specialist to examine the alert in time.

Every alert should state why it was issued, for example: "Water use in this sector has risen above its usual range over the past six hours, while no similar change has appeared in neighbouring sectors." This is clearer and more useful than a generic message saying: "An anomaly was detected." Third: Semi-supervised learning Semi-supervised learning is used when we have a limited number of examples that a specialist has reviewed and assigned the correct outcome to, alongside a larger number of images or records that no one has yet reviewed. Suppose a farm has thousands of images of plant leaves, but a plant pathologist has been able to review only a few hundred of them. This smaller set contains known cases, such as:

  • A healthy leaf.
  • Symptoms of a specific disease.
  • Symptoms of a nutrient deficiency.
  • An image that is insufficient for a judgement.

The remaining images are available, but their status has not been documented. The system first learns from the examples reviewed by the specialist, then makes use of the similarities and differences found in the other images. In this way, a large quantity of data can be used without requiring the expert to review every image manually. But reducing manual work does not mean dispensing with human verification, and it does not mean that a large number of unreviewed images can fix a weak or narrow documented set. If all the images reviewed by the specialist came from:

  • One farm.
  • One cultivar.
  • One phone.
  • Similar lighting conditions.
  • An advanced stage of the disease.

then the system may learn that these limited conditions represent all cases. It may work well on similar images, then perform poorly when it encounters early symptoms, another cultivar, a different farm, or a lower-quality camera. A system that learned what the disease looks like at an advanced stage may fail to recognise its early onset. A system trained on one cultivar may confuse the natural characteristics of another cultivar with disease symptoms. And thousands of images that no specialist has reviewed cannot automatically compensate for the narrowness of the documented examples from which the system learned what the cases mean. We should therefore ask:

  • Do the documented examples represent different farms, cultivars, and stages?
  • Who determined the status of each example, and how was it verified?
  • Are there indeterminate cases, rather than forcing every example into a definite outcome?
  • Did the specialist review a sample of the outcomes later proposed by the system?
  • Did the weakness appear in a particular category, location, or growth stage?
  • Did the result actually improve, or did the system merely become more confident?

Fourth: Learning general features before defining the task There are also methods that first learn recurring general features from a large quantity of images or signals before the system is trained on a specific agricultural task. The system may learn to distinguish shapes, edges, textures, and shades of colour without every image being linked to the name of a disease or an agricultural condition. After that, a smaller set reviewed by a specialist is used to teach the system the required task, such as distinguishing between healthy leaves and leaves showing particular symptoms. This step can reduce the amount of work required, but it does not mean that the system understood the disease or the crop on its own. It learned general visual or numerical features, and was then directed towards a specific agricultural task. The quality of the result remains tied to the soundness of the documented examples and to testing under independent field conditions. Practical example: Dividing the field into three zones Suppose the system analysed moisture data, yield data, and aerial images, and then divided the field into three zones with different characteristics. This result does not directly mean that the field needs three fertiliser prescriptions. The correct next step is to use the map to guide the inspection:

  1. The team first confirms the integrity of the sensors and the data.
  2. The team visits representative points in each zone.
  3. It examines the soil, drainage, shade, and management history.
  4. It takes appropriate samples when needed.
  5. It compares the computational interpretation with field observation and measurement.
  6. It then determines whether the zones really do require different treatments.

The examination may conclude that one zone differs because of the soil, that the second has a drainage problem, and that the third appeared different because of a sensor error. If the three groups were then converted directly into three fertilisation prescriptions, different causes would have been treated as though they were a single problem. Thus, in this case the algorithm provides a map of the questions that merit examination, not a ready-made prescription for answering them. Conclusion When documented outcomes are few, these methods can make use of abundant data in different ways:

  • Unsupervised learning groups similar cases together, but it does not establish why they are similar.
  • Anomaly detection alerts us to what has departed from the pattern, but it does not provide a final diagnosis.
  • Semi-supervised learning combines examples reviewed by a specialist with more data that have not been reviewed, but it may transfer the shortcomings of the limited examples to the rest of the system.
  • Learning general features reduces the need to review every image from the outset, but it does not eliminate field testing or expert judgement.

And the unifying principle is: Abundant data help the system to see patterns, but agricultural meaning does not come from quantity alone; rather, it comes from sound measurement, documented examples, field context, and human verification.

4. Computer Vision: What Does the System See in the Image?

When a farmer or specialist looks at a plant leaf, they do not see colour and shape alone. Rather, they relate what they see to the plant's age, the spread of symptoms, the condition of the field, the weather, irrigation, and their experience with previous cases. A computer vision system, by contrast, begins from something narrower: an image or video clip converted into digital data that a computer can analyse. Computer vision is a set of methods that help the system find patterns within images and video. It may use models called convolutional neural networks, or newer vision models, to recognise shapes, colours, textures, and spatial relationships. Its agricultural applications include inspecting plant diseases, identifying weeds and fruits, estimating ripeness, sorting products, and monitoring animal movement and behaviour. Scientific reviews survey the breadth of these applications in agricultural and food engineering [SRC001][SRC002]. But saying that the system 'understands the image' may suggest that it sees and interprets it as a human does. In fact, the same image can be used to answer different questions, and each question has its own method, result, and limits. 1\. Determining the image's overall state — classification In this task, the system looks at the image as a single unit, then selects the closest state from a set of states it has previously learned. The result might, for example, be:

  • The leaf appears healthy.
  • There are signs suggestive of infection.
  • The fruit appears to fall within a certain quality grade.
  • The image is not sufficient to yield a reliable result.

This type of analysis is useful for initial sorting. But it does not necessarily show where the sign is within the image, does not establish its cause, and does not turn probability into a final diagnosis. So if the system says that the image is 'suspicious', this means that the general pattern resembles examples it has learned from. That alone does not mean that the disease has been confirmed, or that the treatment is now known. 2\. Determining the location of the object or sign — detection Sometimes it is not enough to know that the image contains a fruit, a weed, or a suspicious spot; we also need to know where it is. In this case, the system draws a visible circle or box around the part it has detected. When examining a plant leaf, for example, a box may appear around a particular spot. When analysing an image of a field, the system may place a box around every fruit or weed it has been able to find. This result helps to answer the question: Where is the thing that the system detected? It also helps the user to verify that the system focused on the correct part. The system may think it has detected a disease symptom when in fact it has focused on a shadow, a patch of soil, or part of the image background. But identifying the location of the spot does not establish the cause of its appearance. The system may correctly identify the location of the damage, while determining the disease or stress that caused it still requires other information. 3\. Drawing the boundaries of the affected area — segmentation We may need to know more than the object's location; we may want to know precisely how much area it occupies. In this task, the system colours the area it believes to be damaged, or draws its boundaries on the image. Instead of placing a broad box around the whole leaf, it tries to identify the affected part within it. This can help with:

  • Estimating the proportion of the leaf that is affected.
  • Measuring the amount of vegetation cover in an aerial image.
  • Separating the fruit from the background.
  • Determining the area covered by weeds within part of the field.
  • Comparing changes in the damaged area over time.

Engineers use the word 'pixels' to describe the small points that make up an image, but the reader does not need this term in order to understand the result. What matters in practice is that the system is trying to draw the precise boundaries of the area it has detected. Even so, determining the area of the damage does not establish which disease caused it. The boundaries may be precise, while the interpretation of the cause is wrong. 4\. Following the object or condition over time — tracking When analysing a video or a sequence of images, the system may need to determine whether the object visible now is the same one that appeared a moment earlier or in a previous image. This can be used in:

  • Tracking an animal's movement inside the pen.
  • Tracking a fruit on the sorting line.
  • Counting objects without counting the same object more than once.
  • Monitoring the spread of a damaged area across successive images.
  • Tracking the movement of a machine or a worker within an operating area.

Here, detecting the object in each image separately is not enough; the system must link its appearances across successive frames. The system may face difficulty when the object disappears behind another object, the lighting changes, animals overlap, or the camera moves. Therefore, a loss of tracking should not be turned directly into the conclusion that the animal, fruit, or machine has actually left the area. 5\. Estimating Distance and Shape — Depth Estimation A flat image shows height and width, but it does not always provide the true distance directly. A robot or sorting machine may need to know how close the object is and what shape it has in space. The system may use one or more cameras, or a suitable sensor, to estimate:

  • The distance between the tool and the fruit.
  • The height of the plant or the size of a particular object.
  • How close a person or animal is to the movement area.
  • The position of the branch relative to the robotic arm.
  • The shape of the object that must be grasped or avoided.

Depth estimation can help the machine approach, but it is not sufficient on its own to guarantee safe movement. The system still needs to know the tool's limits and speed, to stop when an obstacle appears, and to verify that the measurement has not changed because of dust, vibration, or poor visibility. Why must the task be defined precisely? The choice of vision task is tied to the action required after the result appears. If the goal is to know whether the image contains fruit, image classification may be enough. But if the goal is to count the fruit, the system must determine each fruit individually. And if the goal is to guide a robotic arm to pick a fruit, it must know its position, distance, and shape, and then convert that information into safe movement. Likewise, knowing that weeds are present in an image is not enough to guide a mechanical tool towards them. The system needs to:

  1. Identify each weed.
  2. Distinguish them from the crop.
  3. Determine its position relative to the machine.
  4. Calculate the tool's movement.
  5. Verify that no person, animal, or unexpected object is present.
  6. Halt execution if confidence falls or the data conflict.

Every transition in this chain adds a new possibility of error. For that reason, it is not valid to describe a system as safe merely because the vision model performed well on test images. Why does a field image differ from a laboratory image? Agricultural images are inherently variable. The same leaf may look different because of:

  • Morning or midday light.
  • Shadow.
  • The camera angle.
  • The type of phone or camera.
  • The lens being closer or farther away.
  • Dust or moisture.
  • The background.
  • The plant's growth stage.
  • Overlapping leaves.
  • The severity of the symptoms and their stage.

Symptoms from different causes may also resemble one another. Some signs of nutrient deficiency may look similar to symptoms of disease or water stress in a single image. Therefore, a system's success in recognising a visual pattern does not mean that it has established the agricultural cause. A model trained on individual leaves placed against a clean background may appear robust in the laboratory, then encounter a very different reality when it sees overlapping leaves, shifting shadows, dust, insects, and symptoms at an early stage in the field. Therefore, the training and test images should represent the conditions in which the system will actually operate, not the best conditions under which an image can be captured. When the system remembers the test instead of proving its capability The images used to train the system must be kept separate from the images used to test it. But that separation may look correct in the files while still being misleading in reality. Suppose we captured ten very similar images of the same plant within minutes, then placed some of those images in the training data and others in the test data. The system may recognise details of the plant, the background, or the camera angle, because the test image is extremely similar to what it has seen before. This is called unintended overlap between the training and test data, and is known technically as "data leakage". It is like training a student on a set of problems, then testing them on the same problems after changing their order or some of their wording. The student may receive a high score, but the test does not prove that they can solve a new problem. Therefore, agricultural testing is stronger when it separates:

  • entire plants, not just images.
  • entire fields.
  • different farms.
  • independent seasons.
  • different imaging devices.
  • varieties or growth stages that were not included in training.

And the real question is not: Did the system succeed on images it had not previously seen as files? Rather: Did it succeed on a farm, in a season, or with a device whose data it had not previously learned from? Why might "overall accuracy" be misleading? A single overall result is not enough to judge a vision model. Suppose, in a simplified numerical example, that the test set contains one thousand images:

  • 950 healthy images.
  • 50 images affected by a significant disease.

If a weak model says that all the images are healthy, it will be correct on 950 out of one thousand images, and its overall accuracy may appear to be 95%. But this model did not detect a single diseased case, and is therefore useless for the task it was created to perform. For this reason, we must ask:

  • How many diseased cases was it able to detect?
  • How many diseased cases did it classify as healthy?
  • How many false alarms did it issue?
  • Does performance differ across diseases or varieties?
  • Can it abstain when the image is unclear?
  • What happens when it encounters a case that was not included in the training data?

If the most dangerous error is to classify a rare and serious case as healthy, then the system's ability to detect that specific case must be measured, rather than settling for an average inflated by common classes. From seeing the object to moving the machine In robotics, the task does not end when the system recognises a fruit or a weed. The result must pass through another chain: Seeing the object ← determining its location ← converting that location into coordinates the machine can understand ← planning the movement ← executing it ← verifying the result ← stopping when danger is present Each part of this chain may be an independent source of error. The model may recognise the fruit correctly, yet the arm may move to an imprecise position because of machine vibration. Or the position may be correct, but a worker's hand may appear within the movement area. Or the connection may be interrupted after the command is sent and before its execution is confirmed. Therefore, the system must include:

  • limits on speed and force.
  • a safety zone around humans and animals.
  • a local stopping mechanism that does not depend on communication.
  • verification that the action succeeded.
  • the ability to return to a safe state.
  • clear authority for the worker to stop or override operation.

And the principle that should remain clear is: A model may succeed in seeing the object while the system fails to handle it safely. Image accuracy is not decision accuracy, and decision accuracy is not execution safety. Practical questions before trusting an agricultural vision system A farmer or buyer can ask:

  • What is the specific task: image classification, localisation, area measurement, or machine guidance?
  • Were the test images collected from independent fields and seasons?
  • Do the images include the lighting conditions, devices, and varieties found on my farm?
  • Which cases does the system fail to distinguish?
  • Does it show examples of errors, or only the best cases?
  • Can it say that the image is insufficient?
  • What happens after the result: an alert, a recommendation, or automated execution?
  • Who reviews the result before a sensitive action is carried out?
  • How does the machine stop if a person appears or confidence drops?

Conclusion Computer vision does not give a machine a human eye or a complete agricultural understanding. It helps the system extract specific information from images and video: identifying a general condition, finding an object, drawing the boundaries of an area, tracking movement, or estimating distance. The value of this information is determined by three things:

  1. Is the visual question framed in a way that suits the decision required?
  2. Do the images represent real field conditions?
  3. Has the result been translated into action through a safe system with clear responsibility?

This enables the reader to distinguish between a system that presents a convincing image in a short experiment and one that has actually been tested to operate under the field's light, dust, diversity, and safety limits.

5. Probabilistic Models and Causal Inference: What Is Expected, and What Can Intervention Change?

Agricultural systems face two different kinds of questions, and their wording may seem similar even though answering them requires different evidence. The first question is: What is likely to happen if current conditions continue? For example, we may ask:

  • What is the expected yield at the end of the season?
  • What is the probability of water stress arising in the coming days?
  • What is the probability that the symptoms are associated with a particular disease?
  • How much production may reach storage in the coming week?

These are predictive questions. Here, the system tries to use what it knows about weather, soil, the crop, and prior management to estimate an outcome that has not yet occurred.

The second question, by contrast, is: What will happen if we change something? For example, we may ask:

  • What will happen to the outcome if the timing of irrigation changes?
  • Will changing one of the management practices improve yield?
  • Will reducing a pump's operating time reduce water use without causing stress to the crop?
  • Was the treatment that was applied the cause of the improvement, or did the improvement occur for other reasons?

These are causal questions, because they do not merely describe what may happen; they ask about the effect of an action we want to take. Correlation does not prove that one factor causes the other. The data may reveal that fields that used a larger quantity of a particular agricultural input achieved higher yields. This is a relationship worth studying, but it does not prove that increasing that input caused the higher yield.

Other explanations may exist, such as:

  • that those fields had better soil from the outset.
  • that irrigation there was more consistent.
  • that the farmers managing them were more experienced.
  • that their owners had better equipment.
  • that the input was used in fields with different crops or varieties.
  • that weather conditions were more favourable at their locations.
  • that the farmer increased the input in response to a condition that was not recorded in the data.

In this case, the input and the yield may appear together, while the real cause of the difference is another factor or a combination of factors. This is similar to observing that farms that use more sensors achieve better results. That alone does not prove that buying sensors will improve the outcome on every farm; farms that can afford to buy sensors may already have been better equipped, better financed, and better managed from the outset. So correlation is useful for discovering patterns and building predictions, but on its own it does not answer the question:

Will changing this factor on another farm lead to the same result? Why does this distinction matter to the farmer? This distinction is not an abstract statistical debate far removed from the field. If the system finds a relationship between a variable and yield, then turns it directly into a recommendation, it has moved from observing correlation to assuming causation. A system that predicts that a given field will achieve a high yield does not necessarily know how to make another field achieve the same yield. Its predictive ability may be based on factors that the farmer cannot change, such as soil type, site history, or seasonal conditions. Likewise, a system may be able to predict which fields are most vulnerable to stress, but that does not automatically prove that any intervention it proposes will prevent stress or deliver an economic benefit. For this reason, three stages must be distinguished:

  1. Observation: What relationship appeared in the data?
  2. Interpretation: What are the possible reasons for this relationship?
  3. Intervention: What evidence is there that changing a particular factor will change the outcome?

The further we move from one stage to the next, the stronger the evidence we need. How do we get closer to knowing the effect of an intervention? To determine whether the intervention was what made the difference, we need a method that separates its effect from the effects of other factors as far as possible. This may be achieved through:

  • an appropriate field experiment comparing specific treatments.
  • allocating treatments in a way that reduces pre-existing differences between the groups.
  • comparing fields or plots that are as similar as possible.
  • measuring conditions before and after the intervention, with a clear baseline.
  • using a causal analytical design that states its assumptions and limitations.
  • drawing on agricultural or biological knowledge that explains how the factor could lead to the outcome.
  • reproducing the result across different sites or seasons.

Nor does using a method called 'causal inference' mean that a relationship has become causal merely because the software has assigned it a number. Every causal analysis depends on assumptions: which factors were measured? Which factors were missing? Were the groups comparable? And did the cause precede the outcome in time? And an algorithm cannot correct for an important factor that was not measured or recorded except by relying on additional assumptions that should be made explicit. The practical rule: a model that is good at prediction is not necessarily good at estimating the effect of an intervention. Accurate outcome prediction does not prove that the system knows how to change that outcome. What do probabilistic models add? In many agricultural situations, the data are not sufficient to give a definitive answer. The picture may be unclear, the symptoms may overlap, the weather record may be incomplete, or the sensor reading may be unstable. Rather than concealing this shortfall, probabilistic models try to express the degree of uncertainty. Let us imagine spots appearing on the leaves of a crop. The possible causes may be:

  • a fungal disease.
  • a nutrient deficiency.
  • water stress.
  • heat or sun damage.
  • mechanical damage.
  • more than one cause at the same time.

If the system relies on a single image, it may not be able to distinguish confidently among these causes. It should not therefore jump straight to a single diagnosis; rather, it can rank the possibilities and indicate what it needs in order to narrow them down. This is what is meant by differential diagnosis: comparing several possible explanations for the symptoms, then gathering information that helps rule some out and strengthen others. For example, the system may ask for:

  • an image of the underside of the leaf.
  • a description of how the symptoms are distributed within the field.
  • whether the symptoms appeared on newer or older leaves.
  • a record of temperature and humidity.
  • information about irrigation or previous treatments.
  • a field inspection or an appropriate test when needed.

As each new piece of information arrives, the system revises the ranking of the possibilities. The value here lies not in the speed of assigning a name, but in organising what we know and what we do not know, and identifying the next most useful piece of information. Probability is not persuasive certainty in numerical form The system may say: The probability of this condition is 80%. But what does this number mean? It does not mean that the individual leaf is '80% infected', nor does it mean that 80% of the leaf is infected. What it means, if the system is well calibrated, is that when it assigns a probability close to 80% to a large number of similar cases, the outcome should occur in about eighty cases out of every hundred, not in all cases. The agreement between confidence numbers and what actually happens is called probability calibration. Suppose a system assigned one hundred cases a probability close to 80%:

  • if the outcome is confirmed in about eighty cases, its confidence is relatively consistent with reality.
  • if it is confirmed in only fifty cases, the system is more confident than it should be.
  • if it is confirmed in ninety-five cases, the system may be less confident than its results justify.

This comparison is not carried out on a single case, but on an appropriate number of similar cases. Ranking cases does not mean that the probability numbers are correct The system may succeed in ranking cases from most likely to least likely, yet still assign misleading confidence numbers. A case to which it assigned 90% may indeed be more likely than one to which it assigned 60%, and in that sense the system's ranking is useful. But if the cases to which it assigned 90% are confirmed only 65% of the time, then its confidence numbers are overstated. So there are two separate questions:

  1. Does the system rank cases in a useful way?
  2. Do the probability numbers reflect reality to an acceptable degree?

It may succeed in the first and fail in the second. Why does confidence change when moving to a new farm? A system's probabilities may be well calibrated in the environment in which it was tested, then become misleading when it is moved to another place. This may happen because of differences in:

  • the cultivar.
  • the growth stage.
  • the climate.
  • the extent of the condition's spread.
  • the camera or sensor.
  • image quality.
  • management practices.
  • The proportion of rare and common cases.
  • The method used to confirm the diagnosis.

Therefore, it is not enough for the supplier to show that the probabilities were appropriate in the development data. Confidence must be monitored after transfer, compared with actual outcomes, and recalibrated when needed. Recalibration does not automatically make the model suitable if the new environment is fundamentally different; new data or a review of the scope of use itself may be required. How should the result be presented to the user? Ideally, the system should not present a bare number or a definitive answer, but rather display the result in four layers: 1\. The most likely result and the uncertainty For example: There are several possible causes, and the current data are insufficient to confirm any one of them. Or: The estimate falls within a certain range and may change if the weather or management changes. 2\. The evidence that influenced the weighting For example: The likelihood of water stress increased because of low moisture readings and high temperatures, but one sensor reading is missing. 3\. The missing information and the next step For example: The sensor needs to be checked, or an additional image is needed, or a description of the distribution of the symptoms, or a specialist review, before the possibilities can be narrowed. 4\. The limits of the decision and responsibility For example: This result is a preliminary alert, not a confirmed diagnosis or an instruction to carry out a treatment. In this way, probability becomes a tool for structuring a decision, not a number that grants the system authority it does not possess. Practical example: Yield prediction is not a recipe for increasing it Suppose a model predicts yield on the basis of:

  • Soil type.
  • Weather.
  • Variety.
  • The timing of operations.
  • Irrigation records.
  • Data from previous seasons.

The model may make a good prediction because it has learned that some fields with better soil and more consistent management produce higher yields. But if it also notices that higher yield was associated with an increase in one of the inputs, it may not infer directly that increasing this input in all fields will raise yield. That increase may be appropriate only for certain fields or conditions, or it may simply be a marker associated with better-managed farms. A predictive model answers the question: What is the likely yield under these conditions? Estimating the effect of an intervention, by contrast, asks: How would yield have differed if we had changed this factor alone, while the other factors remained comparable? The second question is harder, because it concerns an alternative outcome that we did not observe in the same field at the same time. That is why it requires design, experimentation, and clearer assumptions than merely observing the relationship in the records. Practical questions before accepting a recommendation based on a statistical relationship The reader may ask:

  • Does the system provide a prediction, or does it claim to estimate the effect of an intervention?
  • What other factors might explain the apparent relationship?
  • Was the intervention itself tested, or was it merely observed in earlier data?
  • Were the fields or groups comparable?
  • What important factors were not measured?
  • Was the result replicated across different locations and seasons?
  • Does the system present a range and a probability, or a definitive number?
  • Was the validity of the confidence figures tested after transfer to the local environment?
  • What further information could reduce uncertainty?
  • Who has the authority to turn the result into action?

The Distinguishing Rule The difference can be summarised in three questions:

  • What is happening now? A measurement or diagnostic question.
  • What is likely to happen later? A prediction question.
  • What will change if we intervene? A causal question that requires stronger evidence.

A responsible system states the type of question it answers, and does not turn correlation into causation, or a forecast into a prescription, or probability into certainty.

6. Optimisation and Operations Research

Prediction says what may happen, whereas optimisation looks for an action or plan among many alternatives under specified constraints. It may allocate limited water among sectors, arrange a collection route, coordinate harvesting, transport, and cooling, or assign workers and machines across competing time windows. An optimisation problem consists of three things:

  1. Decisions that can be changed: When should we irrigate? Which route should the vehicle take? Which batch should be cooled first?
  2. Constraints that must not be ignored: pump capacity, time, stock, safety, energy, labour, and agricultural boundaries.
  3. An objective we want to improve: reducing water, time, or losses, or balancing several objectives.

Let us imagine three sectors that need irrigation, while pumping capacity does not allow them to operate at the same time. The system may propose a schedule that takes account of water status, crop priority, and energy prices. But if the objective function minimises electricity consumption alone, it may delay a sensitive sector. And if it maximises expected yield alone, it may ignore a small farmer's share or a limit on water depletion. The algorithm executes what has been defined for it; it does not, of its own accord, add the values the designer forgot to specify. For this reason, there is no 'optimal solution' in the abstract. There is an optimal solution with respect to an objective, constraints, and data at a particular moment. If priorities change, a resource fails, or confidence in the prediction declines, the system should recalculate or switch to an alternative plan. In real-world problems, objectives are multiple. We want to preserve yield, reduce water and energy use, avoid stress, and respect the work window. These values cannot always be collapsed into a neutral number. Therefore, presenting several scenarios and their trade-offs is more honest than presenting a single plan as though it were the only answer. The system should also have a safe manual mode. If communication is lost, weather data fail, or constraints conflict, the core process must not stop or continue on an old plan without warning. The power of optimisation is measured not only by the elegance of the plan under ideal conditions, but by its ability to deal with real constraints and exceptions. Practical question: When a vendor says, 'we optimise irrigation', ask: what objective is the system optimising? Who set the weights for water, yield, and energy? And what is the plan if the data are incomplete?

7. Reinforcement Learning and Control: How Does the System Learn from a Sequence of Actions?

In some agricultural problems, the task is not to issue a single forecast and then stop. Every action that is taken changes the state of the farm or the machine, and the new state affects the decision that comes next. When a window is opened in a greenhouse, the temperature may fall and the humidity may change, but the effect of opening also depends on the outside air temperature, the wind, and how long the window remains open. And when a robot moves between rows, its position changes and so do the things its cameras can see; a worker, an animal, or an obstacle may appear in front of it that was not there moments earlier. In such problems, we are not dealing with a single decision, but with a sequence of interrelated decisions. This is where reinforcement learning can have a role. What is meant by reinforcement learning? Reinforcement learning is a method in which the system learns how to choose a sequence of actions by observing the results of what it does. The word 'reinforcement' does not mean that this type of learning is better or more powerful than all other methods. What it means is that, after its actions, the system receives numerical signals that reinforce the choices that bring it closer to the goal, and reduce the value of the choices that move it away from that goal or lead it to undesirable outcomes. The idea can be compared to a learner trying several ways of reaching a result, then gradually retaining the ones that led to better outcomes. But this comparison has limits; the system does not understand the farm in ethical or agronomic terms, and it does not know by itself what should count as 'better'. Humans are the ones who determine what it measures, what actions are available, and what limits must not be crossed. Reinforcement learning consists of several main elements:

  1. State: the information that describes the current situation as far as the system can measure it.
  2. Action: the option the system can execute at that moment.
  3. Outcome: what changed in the environment after the action was carried out.
  4. Numerical score: a signal indicating whether the outcome moved closer to the goal or further away from it.
  5. Selection strategy: the method the system learns for choosing the appropriate action in each state.

Specialists call the program that chooses actions the agent, the numerical score the reward, and the selection strategy the policy. But understanding the functions matters more than memorising the names. A simplified example from a greenhouse Let us suppose there is a system that helps manage temperature, humidity, and energy inside a greenhouse. The state the system sees The available information may include:

  • The temperature inside the greenhouse.
  • Humidity.
  • The outside temperature.
  • The intensity of sunlight.
  • Wind speed.
  • The status of the windows and fans.
  • The moisture of the growing medium.
  • The time of day.
  • The condition of the sensors and equipment.

This is not the "whole environment", but the part that the devices can measure and send to the system. So if one of the sensors is faulty, or if the data arrive late, the system may make its decision on the basis of an incomplete or outdated picture. The actions it can choose The available actions may include:

  • Opening a window to a specified degree.
  • Closing it.
  • Turning on a fan.
  • Reducing or increasing the fan speed.
  • Deploying the shade.
  • Maintaining the current state.
  • Requesting worker intervention instead of automatic execution.

Nor should the system be allowed every action one can imagine. The available actions must be defined in advance and constrained by the limits of the equipment, safety, and the crop. What happens after the action? If the system opens the window, the temperature may fall, but the humidity may also change. Energy consumption may increase as well if the system compensates for the change by operating other equipment. And an action may be suitable at one time of day and unsuitable at another. The system observes what happened after the action, then uses the result to improve its later choices. So it does not learn from a single decision, but from a sequence: State → action → change in the environment → new state → new action What does "reward" mean? Reward in this context is not a prize, nor does it mean that the system feels satisfaction or punishment. It is a numerical score defined by the designer to evaluate the outcome of an action. The score may rise when:

  • Temperature and humidity remain within the required range.
  • Energy consumption falls without harming plant conditions.
  • The system avoids operating equipment unnecessarily.
  • It preserves environmental stability rather than causing sharp changes.
  • It requests worker intervention when data are insufficient or conflicting.

And the score may fall when:

  • Conditions go beyond acceptable limits.
  • The system uses more energy or water than necessary.
  • It switches equipment on and off repeatedly in a way that accelerates wear.
  • It ignores a faulty sensor.
  • It chooses an action that comes close to safety limits.

So the phrase "a reward that expresses the goal" really means: An equation or calculation method set by the designer to convert several outcomes — such as plant condition, energy, water, and equipment stability — into signals that help the system compare its actions. But this score does not necessarily represent the whole goal. What the designer does not include in the calculation is something the system may not take into account. When the system hits the number but misses the purpose Suppose the designer gave the system only one goal: minimising energy consumption. The system may learn to reduce the operation of the fans or ventilation more than it should. It may succeed, in computational terms, in cutting electricity use while the temperature or humidity inside the greenhouse deteriorates. Suppose the goal is to hold the temperature at a single value. The system may switch the heating and cooling equipment on and off repeatedly, ignore humidity, or consume large amounts of energy in order to prevent a small natural variation that does not harm the plant. In both cases, the system has computationally done what it was asked to do, but it has not achieved the broader agricultural purpose. The error here does not necessarily stem from a weakness in the algorithm. The error may lie in the way humans formulated the objective. The system can optimise the score it was given, but it does not automatically know the interests and constraints that were not included in its definition. Therefore, depending on the problem, the evaluation method must include:

  • Plant conditions.
  • Water and energy.
  • Equipment stability.
  • Human and animal safety.
  • Operating limits.
  • The cost of frequent changes.
  • The ability to fall back to a safe mode.
  • Cases in which the decision must be handed over to a human.

The basic rule: the system does not learn what we want in our minds; it learns what we expressed in the data, scores, and constraints. Why is a computational penalty not enough to protect safety? A designer may think of lowering the score when the system carries out a dangerous action. But some actions must not be allowed to be tried in the first place, even if they would later lead to a negative score. For example, the system must not be allowed to learn the danger of shutting off ventilation in severe conditions by carrying out the action and putting the plant at risk. Nor should a robot learn that approaching a worker is unsafe by colliding with that worker once. Therefore, a distinction must be made between:

  • Objectives that the system can optimise.
  • Safety limits that it must not cross.

The system may try to reduce energy use or improve stability within the permitted range, but an independent layer must prevent it from:

  • Exceeding a dangerous temperature limit.
  • Operating a tool when a human is in the movement zone.
  • Continuing if critical sensors fail.
  • Executing a command with a unit or value outside the approved limits.
  • Continuing to move after losing location or communication.
  • Disabling the manual stop mechanism.

These limits are sometimes called hard constraints, meaning binding operational rules that the system cannot break in pursuit of a better score. What is meant by exploration? For the system to know which actions are best, it may need to try options that it has not used much before. This is called exploration. But exploration in a real system differs from experimentation inside a digital game. An error in the field may waste water, stress a crop, damage equipment, or put a worker at risk. Therefore, the system should not be free to try any action on the farm. Risk can be reduced through a gradual path:

  1. Learning or evaluation begins in a simulation that represents the environment reasonably well.
  2. Historical data are used to understand what happened under previous conditions, while bearing in mind that they do not cover all possible actions.
  3. The system is tested in monitoring mode, proposing actions without carrying them out.
  4. Its proposals are compared with what the operator did and with what happened afterwards.
  5. It moves to a limited trial within a low-risk range.
  6. The independent constraints and human stop authority remain active.
  7. Operation is expanded only after sufficient evidence of safety and stability has emerged.

Simulation alone is not enough, because it is a simplified picture of reality. The response of the plant, the equipment, or the weather may differ from what the model assumed. This is sometimes known as the gap between simulation and reality. Therefore, the transition to real operation must be gradual, monitored, and reversible. Not every control problem requires reinforcement learning Reinforcement learning may seem attractive because it learns from a sequence of decisions, but it is not automatically the best option. If the operating rules are clear and stable, a straightforward rule-based system may be more transparent and easier to test. If the equipment response is known and can be represented well, conventional control methods or constrained optimisation may be safer and less costly. Reinforcement learning becomes a candidate when:

  • decisions are sequential, and each action affects what follows.
  • the outcome of some actions appears only after a delay.
  • there is a trade-off between an immediate benefit and a later effect.
  • it is difficult to write a fixed rule that covers all cases.
  • a suitable training environment or simulation is available.
  • a safe range of actions can be specified.
  • there is a way to monitor performance and revert.

But if the goal is merely to detect a disease in an image, predict yield, or issue a single alert, other methods may be more suitable. Reinforcement learning should not be used merely because it sounds newer or more exciting. Who does what within the control system? In a real agricultural system, several functions must be kept separate: 1\. State measurement Sensors and cameras collect information about the environment and the equipment. This information may be incomplete, delayed, or inaccurate. 2\. State estimation or prediction A model may be used to estimate what is not measured directly, or to predict what may happen after a particular action is taken. 3\. Action selection The control strategy determines the proposed action in light of the state, the objective, and the constraints. 4\. Safety check An independent layer reviews whether the action is permitted, whether the data are sufficient, and whether the equipment and environment are in a safe condition. 5\. Execution Commands are sent to the windows, fans, pumps, or motors, and it must be confirmed that the command was actually carried out, not merely sent. 6\. Outcome verification The system compares what was expected with what happened after execution. If the device does not respond, or moves differently, this must be detected. 7\. Human oversight A human has the authority to approve, stop, or override, and can return to clear manual operation when the application or the connection fails. This separation is important because the success of one part does not prove the success of the whole system. The software may select an appropriate action, but the valve does not respond. The valve may execute the command, but the sensor that confirms the outcome may be faulty. The components may function technically, but the objective itself may be incomplete. What is meant by deep reinforcement learning? Deep reinforcement learning is reinforcement learning that uses deep neural networks to help the system represent complex states, estimate the value of actions, or select them. And the word 'deep' does not mean that the system:

  • understands the environment in a human way.
  • is safer.
  • necessarily requires a robot.
  • automatically sees images.
  • is suitable for field deployment.
  • is better than rules or conventional control.

Deep reinforcement learning may be used in a simulated environment to select actions, while computer vision uses another model to analyse images. The two may operate within a single robot, but they perform two different functions. An important source correction: seeing the weed is not controlling its removal In earlier editorial material for this book, a phrase appeared describing one of the weed-detection studies as an example of 'deep reinforcement learning' used in robotic removal. But the source that could be verified \\(SRC009\\) concerns the classification of images captured in soybean fields using convolutional neural networks. That is, the study supports one of the tasks of computer vision: analysing the image and distinguishing between patterns. And this source does not establish that the study developed:

  • a system that learns a sequence of actions.
  • a control strategy in a robot.
  • a reward signal for evaluating removal.
  • automated movement towards the weed.
  • field deployment for removal.
  • a system based on deep reinforcement learning.

The difference is fundamental: Computer vision asks: is there a weed in the image, and where is it? Reinforcement learning and control, by contrast, ask: What action should the machine take now, and how will it affect the state and the next action? A weeding robot may need a vision model to identify the plant, then another system to determine the location, then a motion planner, then a controller to execute the action, then a safety layer that stops the machine when danger arises. The success of a classification algorithm does not prove the success of this chain. Including this correction is not an argument about naming, but an application of the book's evidentiary method: no algorithm, operational capability, or field impact should be attributed to the source unless it establishes it. Integrated example: ventilation management in a greenhouse Suppose the system manages a window and a fan inside a greenhouse. At a given moment, it reads:

  • a slightly elevated internal temperature.
  • high humidity.
  • a lower external temperature.
  • moderate winds.
  • no equipment alarm.

It can choose between:

  • leaving things as they are.
  • opening the window partially.
  • opening it further.
  • switching on the fan.
  • using the window and the fan together.
  • requesting a worker review.

If it opens the window partially and the temperature and humidity fall without high energy consumption, the evaluation method records that the outcome has moved closer to the goal. If the action causes an abrupt change or the humidity moves outside the range, the numerical score declines. But there must still be rules that are not subject to experimentation, such as:

  • not operating the equipment if the maintenance cover is open.
  • not continuing when critical sensors conflict.
  • not exceeding the operating limits of the motors.
  • switching to a safe mode when communication is lost.
  • allowing the worker to stop it immediately.

In this way, the system learns within a permitted space and does not have the freedom to redefine safety through experimentation. Practical questions before accepting a system based on reinforcement learning The farmer, buyer, or operator can ask:

  • What sequential decisions does the system learn to make?
  • What information describes the current state?
  • What happens if an important sensor is faulty or delayed?
  • What actions is the system allowed to take?
  • How was the numerical score that guides learning calculated?
  • Which aspects were not included in this numerical score?
  • What limits can the system not exceed?
  • Where was the training conducted: in simulation or in actual operation?
  • How was the gap between the simulation and the farm addressed?
  • Has the system begun with recommendation and monitoring before execution?
  • What happens when it encounters a situation it has not seen before?
  • Is there a local stop mechanism independent of the application?
  • Who has the authority to approve, override, and return to manual mode?
  • Is reinforcement learning truly necessary, or is there a simpler, clearer approach?

For the advanced reader: Evaluating a reinforcement learning system is not limited to the average score it achieved during training. Its behaviour in rare cases, the stability of its performance, its sensitivity to how the objective is defined, its ability to operate with incomplete data, its observance of constraints, the gap between simulation and reality, and the safety of the transition from recommendation to execution must all be examined. Conclusion Reinforcement learning is not a method that gives the machine the freedom to experiment in the field. It is an approach to learning how to choose sequential actions within limits designed by humans. Its value depends on five things:

  1. That the available information truly describes the relevant state.
  2. That the numerical score represents the agricultural objective without reducing it to a misleading number.
  3. That safety remains in independent rules that must not be overridden.
  4. That the transition from simulation to operation is gradual and reversible.
  5. That human beings remain the ones with the authority to stop the system and give final approval.

And the rule the reader should take away is: The system does not learn the objective we implicitly intend; it learns the behaviour that the numbers reward and the constraints permit. Therefore, the objective and the limits must be tested before the intelligence of the algorithm is tested. The final safety rule can be stated more clearly as follows: The field should not be the first place where the designer discovers that the method for evaluating the system is incomplete, or that trying a new action may harm the crop, people, or equipment.

8. Language Models, Generation, and Retrieval: From the Question to the Grounded Answer

A farmer can ask the digital assistant in ordinary language instead of searching through long menus or filling out a complex technical form. They may ask about the meaning of an alert, request a summary of a document, ask for instructions to be translated, compare two records, or extract a specific piece of information from a report. This ease is one of the most important benefits of language models, but it is also a potential source of misunderstanding. An answer written in clear, confident language may seem to come from an expert who knows the field, the location, and the product, when in fact the system may have composed it from linguistic patterns or from incomplete information that is not sufficient for decision-making. So it is not enough to ask: Can the system write an understandable answer? We must also ask: Where did the information in the answer come from? And does it apply to this crop, place, time, and decision? What do language models do? A language model is a system trained on large quantities of text in order to learn the relationships that recur among words, sentences, and ideas. When it is asked for an answer, it builds the text piece by piece on the basis of the question, the system instructions, and whatever context is available to it. It can be used for tasks such as:

  • Summarising a long document.
  • Translating content from one language into another.
  • Turning scattered notes into an organised report.
  • Extracting the crop name, location, date, or unit from a record.
  • Explaining a technical term in simpler language.
  • Comparing several documents.
  • Managing a dialogue consisting of more than one question and answer.
  • Formulating an answer in natural language that the user can read.

But these capabilities do not mean that the model has independent agricultural understanding or an up-to-date memory of every regulation, price, and product. It is good at constructing coherent text, and in doing so it may use correct information, incomplete information, or a connection that appears linguistically plausible but is not firmly grounded in the source. Fluency, then, is a property of the wording of the answer, not evidence of the accuracy of its content. How can an answer be convincing yet unsupported? When the model does not have enough information, it may not leave the gap visible. It may complete the sentence with wording that seems plausible on the basis of patterns it learned from earlier texts. This may result in:

  • A number that does not appear in the source.
  • An uncertain cause framed as though it were a diagnosis.
  • A dose that seems reasonable but is not tied to a registered product.
  • A product name that is not suitable for the crop or the country.
  • An outdated price, or one from another market.
  • A citation that appears official but does not support the claim.
  • A translation that omitted a qualification, changed a unit, or overstated the level of certainty.
  • A general answer presented as though it were a farm-specific recommendation.

This does not mean that the model is deliberately trying to deceive. The problem is that it is designed to produce text that fits the linguistic context, not to guarantee that every sentence has been checked against a source that is valid for this decision. The basic rule: A system's ability to complete a sentence does not mean it has the evidence needed to complete a decision. A short agricultural question may conceal a great deal of information If a farmer asks: "Spots have appeared on my crop; what should I use?" the sentence may seem clear linguistically, but it is agriculturally incomplete. It does not tell us:

  • What is the crop?
  • What is the variety?
  • What is the plant's age and growth stage?
  • Where is the farm located?
  • When did the symptoms appear?
  • Did they appear on young leaves or old ones?
  • Are they on the upper surface or the lower one?
  • Are they spread across the whole field, or confined to limited patches?
  • Are there symptoms on the stem, root, or fruit?
  • What are the weather and irrigation conditions?
  • Has any previous treatment been used?
  • Is the image clear, and is it representative of the case?
  • Which products are registered for the crop and pest in the country of use?
  • Are there restrictions related to the pre-harvest interval or worker re-entry?

If the system fills in these gaps on its own, it may give an answer suited to a case it has imagined, not to the case actually present on the farm. For this reason, a responsible system should stop when information is missing and ask for what it needs before narrowing the possibilities. The best initial answer may be a clarifying question, not a recommendation. What is the difference between generation and retrieval? Generation is the process of formulating a new answer in natural language. The model takes the question and the available context, then organises them into text that the user can read. Retrieval, by contrast, is the search within a collection of documents or records for passages that may be relevant to the question. If the user asks about the required waiting period for a specific product, the system may search in:

  • The current approved product label.
  • The product registration record.
  • Documents issued by the competent regulatory authority.
  • Official guidance related to the crop and pest.
  • A dated edition of an approved knowledge base.

Retrieval does not necessarily formulate the final answer; rather, it finds the material from which the answer can be built. When the system combines searching sources with formulating an answer on the basis of what it has retrieved, this workflow is called retrieval-augmented generation, usually abbreviated to RAG. What happens in retrieval-augmented generation? The workflow can be simplified into six stages:

  1. Understanding the question: determining what the user is asking and what information is missing.
  2. Defining the search scope: choosing the set of sources appropriate to the crop, location, and topic.
  3. Finding candidate passages: searching for parts of documents that appear relevant to the question.
  4. Checking validity: excluding documents that are outdated, unofficial, or unsuitable for the context.
  5. Formulating the answer: using the accepted passages to write a clear answer.
  6. Showing the source and the limits: enabling the user to know where the information comes from and what the answer cannot establish.

This pathway may reduce unsupported answers, but it does not prevent them automatically. The system may err at any of the six stages. Retrieval does not plant permanent knowledge inside the model When the system retrieves a document at the time of the question, it places part of it before the model to help it formulate the answer. That does not mean the knowledge has become a permanent part of the model, or that all its later answers will use it correctly. With a new question, a new search may be run and a different passage retrieved. Therefore, the quality of each answer depends on:

  • the wording of the question.
  • the set of available sources.
  • how they are organised and indexed.
  • the search criteria.
  • the passages presented to the model.
  • the instructions that constrain the wording.
  • the system's ability to abstain when it does not find sufficient support.

This means that evaluating the language model alone is not enough; the entire search, formulation, and citation pathway must be evaluated. How can the retrieval pathway go wrong? The pathway may fail in several ways: 1\. Misunderstanding the question It may interpret a local word as the name of a disease, or confuse the name of a crop with the name of a variety, or assume that the user is asking about treatment when in fact the user is asking about initial identification. 2\. Searching in an unsuitable collection The system may search general articles when the question requires an official local document, or search guidance for a different crop. 3\. Retrieving an old document The document may have been correct when it was issued, but the product registration, restrictions, price, or instructions may have changed afterwards. 4\. Retrieving a valid source for a different context The recommendation may be correct for another country, another variety, another growth stage, or a different production system. 5\. Selecting a passage that shares the words but not the meaning The passage may contain the disease name and the crop, but discuss prevention rather than treatment, or describe a research study rather than operational instructions. 6\. Going beyond the source during formulation The source may mention a general note, and the model may then turn it into a specific recommendation that does not appear in the source. 7\. Placing a correct citation after a sentence that the source does not support The presence of the source name or its link is not enough. The claim itself must appear in the source, with the same degree of certainty and the same limits. So the source may be reliable, and retrieval may be successful at the textual level, yet the answer may still be unfit for agricultural decision-making. Similarity in wording is not fitness for use Suppose the system finds a document discussing leaf spots in a particular crop. The document may appear highly relevant to the question, but it may:

  • concern a country where product registration differs.
  • deal with a different variety or environment.
  • be a laboratory study rather than field guidance.
  • describe similar symptoms with a different cause.
  • be old or have been replaced by a newer version.
  • discuss prevention while the user is asking about an existing case.
  • mention an active ingredient without a locally registered product.
  • be in translation and have lost an important qualification.

Therefore, retrieval should not rely on word similarity alone. The search must be narrowed using information such as:

  • Crop.
  • Variety, if needed.
  • Location and legal jurisdiction.
  • Growth stage.
  • Date.
  • Document type.
  • Issuing body.
  • Document status: in force, archived, draft, or revoked.
  • Product registration status.
  • Decision type: general information, diagnosis, dosage, or regulated action.

These conditions are sometimes called search filters. They are rules that prevent the system from treating every linguistically similar text as relevant to the question. Who approves the sources? The phrase 'approved sources' should not remain vague. We need to know:

  • Who selected the source?
  • What kind of authority does it have?
  • Is it a scientific, regulatory, or commercial source?
  • When was it published or updated?
  • Is it still in force?
  • For which region, crop, and production system is it applicable?
  • Which claims can it support?
  • What may not be inferred from it?

A scientific article may help in understanding a phenomenon or a research finding, but it does not establish a product's legal registration. And the official directions for use set out the conditions for a registered product, but they do not by themselves prove a general economic effect. And a price published in an official market may be correct for that market and date, but it does not necessarily equal what a farmer in another region receives after transport and commission. So the authority of a source varies with the question. Three layers are not always enough It is useful for an answer to distinguish between what the sources say, what the system has inferred, and what requires a human decision. But in sensitive agricultural decisions, it is better to present five layers: 1\. What the system understood from the question For example: I understood that you are asking about the cause of spots that appeared on the leaves of a particular crop at a specific growth stage. This gives the user an opportunity to correct that understanding before the answer is built. 2\. What the sources say The claim supported by the source is presented, together with the source's name, date, and scope. 3\. What the system has assembled from several pieces of information If the system combines more than one source or infers a relationship, it must state that this is its own formulation or comparison, not a direct quotation. 4\. What remains unknown Such as the lack of a clear image, the absence of a test, not knowing the country, or not verifying whether the document is still valid. 5\. What requires a human decision or a competent authority Such as the final diagnosis, the choice of a regulated substance, approval of the dosage, or carrying out an intervention that may harm the crop or people. This prevents supported information from being conflated with automated inference or professional judgement. Citation is not decoration The user should be able to examine the citation and understand its relationship to the answer. A good citation makes clear:

  • The name of the document or issuing body.
  • The date of issue or last update.
  • The part that was used.
  • The scope of the document.
  • Whether it is official, scientific, or commercial.
  • Whether the system reproduced the text, summarised it, or combined it with another source.

Placing a general link after a long paragraph does not show which sentence the source supports. It may give the answer an appearance of documentation without enabling the reader to trace the claim. Rule: It is not enough for the source merely to exist; it must be relevant to the question, it must support the sentence, and the limits of the sentence must remain within the limits of the source. What does the system do when it cannot find sufficient support? Abstention is not a failure of dialogue. The system can continue to help the user without inventing an answer. It might say:

  • I cannot find a valid source for this product in the specified country.
  • The available document is outdated and is not sufficient to establish the dose.
  • The image does not allow the possible causes to be distinguished.
  • The question requires the crop name, growth stage, and location.
  • There is general information, but it is not sufficient for a treatment recommendation.
  • A specialist examination or a test is required before choosing the intervention.

It then offers a safe next step:

  • Request an additional image.
  • Specify the location.
  • Give the product name exactly as it appears on the package.
  • Check the registration status.
  • Review the current official directions for use.
  • Contact the extension service or a qualified specialist.

In this way, abstention becomes not a dead end but a clear path for completing the evidence. Example of an irresponsible course The user asks: "Spots have appeared on my crop; what should I use?" The system answers directly: "Use product X at such-and-such a dose." This answer is irresponsible because it moves from a general description to a substance and a dose without knowing:

  • The crop.
  • The country.
  • The diagnosis.
  • The growth stage.
  • The registration status.
  • The official directions for use.
  • The safety interval.
  • Whether there are other causes of the symptoms.
  • The user's ability to carry out the treatment safely.

The wording may be linguistically sound, yet it lacks a basis for the decision. Example of a responsible answer The answer might be: "The current information is not sufficient to determine the cause of the spots or to choose a treatment. These symptoms may overlap among disease, nutrient deficiency, water stress, and weather-related damage. Give the crop name and variety, if known, the country or region, the growth stage, and where the symptoms appear on the plant, and send a clear image of the upper and lower surfaces of the leaf, as well as an image showing how widely the problem is spread in the field. I can then compare the symptoms with appropriate sources and set out the possibilities and the missing information, but specifying a product or a dose requires verification of the diagnosis, the product's registration for the crop and pest in the country of use, and the current official directions for use, and it may require consulting a specialist." This answer does not merely abstain; rather, it:

  • Explains why the information is insufficient.
  • Shows that more than one interpretation is possible.
  • Requests specific data.
  • Clarifies what the system can do.
  • Identifies what requires an official source or a specialist.
  • Prevents linguistic fluency from turning into unwarranted therapeutic authority.

Another example: a question about price If a farmer asks: "What is the price of my crop today?" it is not enough for the system to retrieve a figure that includes the crop name. It must know:

  • The country, region, and market.
  • The date of the price, and the time if needed.
  • The currency.
  • The unit of sale.
  • The variety.
  • The quality grade.
  • Whether the price is wholesale, retail, or farm-gate.
  • Whether it includes transport or commission.
  • Which entity published it?

And a responsible answer does not say: "The price is such-and-such." Rather, it says, for example: "I found a published price in the specified market on such-and-such a date, for the unit and grade mentioned. The price received by the farmer may differ depending on quality, transport, and commission, and I have not verified a local transaction on your farm." This makes clear that retrieval returned specific market information, not the 'real price' everywhere. Translation itself requires governance A language model may be used to translate a product label or agricultural guidance. But a fluent translation may conceal an error in:

  • The unit of measurement.
  • Negation.
  • The waiting period.
  • The name of the pest.
  • The registered crop.
  • The degree of obligation.
  • The warning.
  • The legal jurisdiction.

Therefore, a machine translation of a safety document should not become an alternative authority to the official text. The system can help with understanding, but the original reference must be retained, its source and version shown, and sensitive terminology reviewed before implementation. How do we evaluate the system? It is not enough to assess stylistic quality or user satisfaction. Every step in the pipeline must be tested: Understanding the question

  • Did the system identify the crop, the location, and the intended meaning?
  • Did it request the missing information?

Retrieval

  • Did it find the appropriate source?
  • Did it exclude an outdated or non-local source?
  • Did it retrieve the part that answers the question?

Wording

  • Did the answer adhere to what the source states?
  • Did it add an unsupported claim?
  • Did it preserve the constraints and uncertainty?

Referencing

  • Can the reader access the source?
  • Does the source actually support the sentence?
  • Are the document's date and status shown?

Safety

  • Did the system refrain from providing a dose or a diagnosis when the evidence was insufficient?
  • Did it make clear when a specialist or regulatory authority should be consulted?
  • Did it prevent general information from being turned into an actionable instruction?

Agricultural application

  • Is the answer valid for the crop, location, and timing?
  • Can the user understand it and carry it out?
  • Did human responsibility remain clear?

The system may succeed in finding a correct document, then fail to interpret it. It may formulate a faithful answer to an unsuitable passage. It may provide a genuine reference for a sentence that the source does not support. Therefore, the whole pipeline must be evaluated, not the model in isolation. For the advanced reader: retrieval reduces the model's reliance on what it previously learned, but it shifts part of the risk to the quality of the index, question understanding, result ranking, document validity, and disciplined generation. For that reason, the pipeline's success is not measured by the mere presence of a source in the answer, but by the soundness of the relationship among the question, the source, the claim, the context, and the decision. The role of Chapter Seven The reference to Chapter Seven may be retained, but it is better to clarify what it will add: Chapter Seven will address in greater detail how to build decision-support and knowledge-retrieval systems, organise sources, constrain search by context, show references, and prevent retrieved text from turning into a recommendation that exceeds its authority. This gives the reference a clear function, rather than leaving it as merely a promise of information to come. Conclusion A language model can make access to knowledge easier: it turns an everyday question into a search, gathers passages, summarises them, and formulates an intelligible answer. But linguistic fluency does not shorten the scientific and regulatory path between question and decision. A responsible system must therefore know:

  • What it understood from the question.
  • What information it lacks.
  • Where it searched.
  • Why it chose the source.
  • Whether the source is current and valid for the context.
  • What it took from the source.
  • What it inferred.
  • Where its role ends.
  • And when it must refrain or escalate the decision to a human being.

The section's message can be summarised in the following statement: A language model can be a gateway to knowledge, but it does not become a source merely because it speaks fluently, nor does it become a decision-maker merely because it found a text close to the question. The proposed closing statement is: A trustworthy answer does not merely appear correct; it reveals its source, the limits of its applicability, what is missing, and who has the authority to turn it into action.

9. Explainable AI

When the system produces an output, a natural question arises: Why did it produce this output? The model may predict a drop in yield, raise the level of suspicion of a disease, issue an alert about water stress, or classify a fruit within a particular quality grade. But the output alone does not tell the user what information the system relied on, nor whether that information was complete or reliable. Explainable AI tools seek to help the user understand the relationship between the data fed into the model and the output it produced. They may show, for example:

  • that low soil moisture was among the factors associated with a higher estimated risk of stress.
  • that the forecast temperature affected the estimated irrigation requirement.
  • that the cultivar and growth stage appeared among the factors associated with the yield prediction.
  • that a particular part of the leaf image was associated with its being classified as suspicious.
  • that the absence of a sensor reading, or a change in it, affected the system's confidence.
  • that the output was close to the threshold separating two classes.

But these tools do not answer every meaning of the word 'why'. They usually explain why the model produced this output according to its calculations, but do not necessarily establish why the phenomenon occurred in the field. Three different questions are hidden inside the word 'why' When the user asks 'why?', they may mean one of three questions: 1\. Why did the model produce this output? This is a question about the model's behaviour. The answer might be: The model raised the risk estimate because the moisture reading fell, the forecast temperature rose, and no recent reading arrived from one of the sensors. This answer shows what information the model used. 2\. Why did the agricultural phenomenon occur? This is a question about agricultural reality. A possible answer might be: Perhaps the moisture fell because water demand increased, or because irrigation was not reaching that sector properly, or because of a soil problem, or because of a sensor error. A model-interpretation tool on its own cannot determine which of these causes is correct. The fact that the model used the moisture reading does not prove the cause of its decline, nor does it prove that the plant is actually under stress. 3\. Why should we take this action? This is a question about the decision. Even if the model's alert is reasonable, it does not automatically follow that the correct action is to increase irrigation. The sensor may be faulty, rain may be expected, the problem may lie in water reaching only a limited area, or the crop stage may differ from what the system assumed. So a distinction must be made between: explaining the model output ← explaining the agricultural situation ← justifying the action And moving from one level to the next requires additional data and evidence. For whom is the explanation intended? There is no single explanation that works for all users. Each party needs a different level. The developer wants to know:

  • Did the model rely on valid data?
  • Did it learn an incorrect shortcut?
  • Did it focus on the image background instead of the plant?
  • Did its behaviour change after the data update?
  • Is the output stable under a slight change in the input?

The agricultural specialist wants to know:

  • Are the factors used by the model agriculturally meaningful?
  • Has it overlooked an important variable?
  • Has it confused a statistical relationship with an agricultural cause?
  • Is the outcome plausible for this cultivar, season, and location?
  • What measurement or test could confirm or refute it?

The farmer or operator needs a practical answer:

  • What did the system say?
  • What information was the outcome based on?
  • Is this information current and reliable?
  • What does the system not know?
  • What is the next safe step?

The reviewer or responsible authority needs an auditable record:

  • What is the model version?
  • What data were available at the time of the decision?
  • What rules or limits were applied?
  • Did a human override the outcome? Why?
  • Was the source or sensor valid at that time?

The visualisation suitable for the researcher or developer may not be useful to the worker who needs to take action within minutes. Therefore, the explanation must be designed for the person and the decision, not only for what the tool can display. Two types of explanation Two main levels can be distinguished: Explanation of a single case It seeks to answer the question: Why did the system give this outcome now, for this image or this field? It may show that soil moisture and temperature were the two factors most strongly associated with a particular alert, or that part of the leaf image was associated with its classification. This explanation is useful for reviewing a specific case, but it does not necessarily describe the model's behaviour in all cases. Explanation of the model's general behaviour It seeks to answer questions such as:

  • What types of information does the model usually rely on?
  • Does the yield prediction increase or decrease as a certain variable changes?
  • Does its behaviour differ between cultivars or regions?
  • Which variables appear frequently among the influential factors?
  • Does it weaken when a certain type of data is absent?

This level helps in understanding the model across a large set of cases, but by itself it does not explain an individual decision. The model may generally rely on weather and soil, while the outcome for a particular field may have been affected mainly by an unusual reading from a single sensor. Therefore, the general explanation should not be used instead of the case explanation, or vice versa. What forms can explanation take? There are multiple ways to present an explanation, and each has its own meaning and limits. 1\. Showing the factors associated with the outcome The system may display a list such as:

  • Low soil moisture: increased the risk estimate.
  • High forecast temperature: increased the estimate.
  • Probability of rain: reduced the estimate.
  • No recent reading: reduced confidence.

This list shows how the model handled the information, but it does not prove that these factors caused the agricultural condition. 2\. Highlighting part of the image An explainability tool may colour parts of the leaf image to show the areas most strongly associated with the result. This may help reveal that the model focused on:

  • The relevant spot.
  • The edge of the leaf.
  • The background of the image.
  • A tag placed next to the plant.
  • A shadow or reflection unrelated to the condition.

But the coloured area does not mean that the model 'saw the disease' as a specialist would, nor does it prove that the highlighted part caused the classification. It is a sign of a computational relationship within the model that requires cautious interpretation. 3\. Comparing the result with a similar case The system may say: This case resembles earlier examples classified in the same category. The comparison helps the user see the basis of the similarity. But the user should know:

  • Are the earlier examples documented?
  • Did they come from a similar crop, cultivar, and environment?
  • Is the similarity in the symptoms, or in the background and lighting?
  • Are there similar examples that led to a different result?

4\. Stating what would have changed the result The tool may say: If the moisture reading had been higher in the model's calculation, the risk estimate would have been lower. This shows the model's sensitivity to a change in the input, but it does not prove that actually increasing moisture would prevent the risk in the field. It describes a change in the model's calculation, not the effect of a real agricultural intervention. Here again, the difference appears between: What changes the model's result? and: What changes agricultural reality? Interpretation may be an approximation, not a full disclosure Some models are relatively simple, and their method of calculation can be traced directly. With some complex models, however, another tool is used after the result is produced to try to approximate the reasons behind it. In the second case, the interpretation is not a complete window onto everything that happened inside the model. It may be a simplification or an approximation of its behaviour. Therefore, the interpretation itself should be tested:

  • Does it change greatly if the data change only slightly?
  • Does it give similar interpretations for contradictory results?
  • Does it truly reflect what the model used?
  • Can it distinguish an important factor from one that is merely associated with it?
  • Did the interpretation remain stable after the model was updated?

So if the list of 'most important factors' changes radically after a slight change that does not alter the result, the interpretation may be unstable, even if the graphic appears convincing. Factors associated with one another The data may contain variables that move together, such as:

  • Temperature and evaporation.
  • The season and day length.
  • Soil type and drainage.
  • Farm size and equipment type.
  • Experience and the quality of record-keeping.

When variables are strongly correlated, the explainability tool may distribute importance among them in different ways. One may appear near the top even though the other carries similar information. Therefore, the ranking of factors should not be read as a fixed ranking of agricultural causes. A specialist may need to examine the relationships among them, and whether the model is using one as a substitute for another. The basic limits of explainability Explainability has limits that must remain visible:

  • It describes what the model used, and does not necessarily establish what caused the phenomenon in the field.
  • It may be an approximation of the model's behaviour, not a complete description of how it computes.
  • It does not turn a wrong model into a correct one.
  • It does not make up for insufficient data or corrupted measurements.
  • It does not prove that the training data represent the new farm.
  • It does not prove that the probabilities are calibrated or that the result is accurate.
  • It does not automatically justify the proposed action.
  • It may seem convincing even when the case lies outside the range of what the model has learned.
  • It may change when a different explainability tool is used.
  • It may hide instability behind a clear, attractive graphic.

Explainability, then, is a tool for scrutiny and accountability, not a certificate of validation. Explainability does not fix data errors Suppose a soil-moisture sensor did not send a recent reading, so the system used the last stored reading from six hours earlier. An explainability tool may show that moisture was an important factor in the result, but it does not make the old reading current. Likewise, if the training images do not represent a particular cultivar or region, the system may explain the result for the new image, but that explanation does not prove that the model is valid for that environment. Explainability should present data quality alongside its effect, such as:

  • The time of the latest reading.
  • The presence of missing values.
  • The sensor's calibration status.
  • The data source.
  • How closely the case resembles the training data.
  • Whether the result falls within the range in which the model was tested.

An influential factor is not useful if its value itself is unreliable. Explainability does not prove causality If the system shows that moisture was an influential factor in the yield prediction, that means the model used moisture information when calculating the result. But it does not prove that changing moisture alone will change yield by the same amount. Moisture may be associated with other factors, such as:

  • Soil type.
  • Irrigation system.
  • Weather.
  • Growth stage.
  • Field management.
  • Cultivar.

To determine the effect of a particular intervention, we need an experiment or a causal design, along with appropriate agricultural knowledge, as the previous section explained. A key rule: model explainability tells us how its calculation changed, not how the field will necessarily change if we alter one of the inputs. Limits of inference from the yield-prediction study The yield-prediction study \\(SRC006\\) provides an example of using regression models with tools that help interpret their results within a specific research setting. The study can be used to understand:

  • The type of data fed into the models.
  • The prediction method used in the study.
  • How the explainability tools were applied in that setting.
  • The results reported by the researchers within the limits of their data and study design.

But the ranking of variables or their importance in the study must not be turned into a general causal rule for every farm. If a particular variable appears influential in the model's prediction, that does not prove:

  • that it is the main cause of yield in all environments.
  • that changing this variable will produce the same effect in another field.
  • that the ranking of factors will remain stable in another variety or season.
  • that the model is suitable for making a direct intervention recommendation.
  • that the computational result implies a general economic or field impact.

The permissible conclusion remains within the limits of the study and the evaluation setting it used. What does the farmer need from explanation? Useful explanation for the farmer does not begin with a technical diagram or a long list of variables. Rather, it answers six questions:

  1. What is the result?
  2. What exactly did the system say?

  3. What information did it rely on?
  4. Which readings, images, or records influenced the result?

  5. What is the condition of this information?
  6. Is it recent, complete, and calibrated, or is there an old or missing reading?

  7. What is the degree of uncertainty?
  8. Is the result clear, close to the threshold, or outside the model's range of experience?

  9. What does the explanation not establish?
  10. Is the result a preliminary alert, a diagnosis, or merely a prediction?

  11. What is the safe next step?
  12. What measurement, inspection, or review is required before changing the decision?

In this way, explanation becomes a working tool, not a technical display detached from the decision. Reframing the example Instead of saying: "The risk of water stress increased because surface-layer moisture fell over three days and the forecast temperature rose..." —a phrasing that may be understood as causal proof—we can say: "The system raised its estimate of the probability of water stress because the surface-layer moisture reading fell over the past three days, and because temperature forecasts increased. But the sensor in sector 4 has not sent a new reading for six hours, so the estimate may be based on incomplete data. Before adjusting the irrigation schedule, inspect the sensor, take an independent measurement from the sector, and compare the result with a neighbouring sector under similar conditions." This phrasing makes clear:

  • that the discussion concerns the model's estimate, not a confirmed cause.
  • which data influenced the estimate.
  • that one data input is incomplete.
  • how that lack affects confidence.
  • what the next step is.
  • that adjusting irrigation has not been turned into an automatic order.

An example from a leaf image The system may say: "The model classified the image as a suspected case, and the result was associated mainly with the spots near the edge of the leaf. But the image does not show the underside, and the model did not confidently distinguish between a possible disease and a nutrient deficiency. Send a clearer image of both sides, and state where the symptoms appear on the plant and how they are distributed in the field before moving to a diagnosis or treatment." This explanation is better than a coloured map alone, because it shows:

  • Where the model focused.
  • Which possibilities overlap.
  • What information is missing.
  • What the next step is.
  • That the classification is not a final diagnosis.

An example from yield prediction The system may say: "The predicted yield came in below this field's historical average. Within the model, the factors most strongly associated with the lower prediction are delayed planting, low soil moisture during a specific period, and higher forecast temperatures. But the record does not contain complete information about the most recent treatment, and the model has not been tested in this field during an extremely hot season. The result should therefore be treated as a planning scenario, not a guaranteed value." This wording prevents the explanation from being turned into:

  • A confirmed cause.
  • A promise of yield.
  • An intervention recommendation.
  • A generalisation beyond the scope of testing.

When is the explanation insufficient? The explanation alone should not be relied on when:

  • Critical data are missing or outdated.
  • The case falls outside the crops or varieties on which the model was tested.
  • The result is close to the decision threshold.
  • Several sensors disagree.
  • The potential error relates to human or animal safety.
  • The decision leads to a dose or a regulated substance.
  • The system controls a moving machine.
  • The user cannot inspect the explanation or its source.
  • The explanation changes radically with a slight change in the data.
  • There is a gap between what the model explains and what the decision needs to establish.

In these cases, the explanation should lead to inspection or escalation, not merely to greater confidence. Practical questions for evaluating the explanation The reader or buyer may ask:

  • Does the system explain the outcome of a single case, or its general behaviour?
  • Does it make clear which data it relied on, and when they are from?
  • Does it show missing or unreliable data?
  • Is the explanation a direct description of the model, or an approximation produced by another tool?
  • Does the explanation change with a slight change in the input?
  • Has the developer examined the possibility that the model is relying on the background or the device?
  • Does the system distinguish between correlation and causation?
  • Does it make clear the range of crops, locations, and seasons in which it was tested?
  • Can the user understand the explanation without programming expertise?
  • Does the explanation lead to a safe, actionable step?
  • Can a human contest the result or override it?
  • Were the data and the version that produced the explanation preserved for later review?

For the advanced reader: the explanation should be evaluated for how faithfully it represents the model's behaviour, how stable it remains under small changes, how understandable it is, and how well it fits the user and the decision. A simpler, clearer model may be better for a sensitive decision than a model that scores slightly higher on a laboratory metric but requires an unstable approximate explanation. Conclusion Explainable AI does not make the model conscious of its reasons, does not turn a statistical relationship into an agricultural cause, and does not automatically confer legitimacy on the decision. Its best role is to help us ask:

  • What information did the model use?
  • Was that information accurate and up to date?
  • Did it focus on what it was supposed to focus on?
  • What did it leave out of account?
  • What are the limits of its confidence?
  • And what should be verified before taking action?

The section's message can be summarised in the following phrase: A good explanation does not merely tell us why the number came out as it did; it reveals the data that produced it, the limits that constrain it, and the step that prevents it from being turned into a blind decision.

10. The Smart Agricultural System Is an Integrated System, Not a Standalone Model

When we say that a farm uses a 'smart irrigation model' or a 'smart disease-detection system', the reader may imagine a single program that receives the data and then issues the decision. But useful agricultural systems generally do not work in this simplified way. They are more like an interconnected chain: it begins by measuring what is happening in the field, passes through checking and analysing the data and comparing it with agricultural constraints, and ends by presenting a recommendation to a human being or by carrying out an automated action within specified limits. For this reason, it is described as a hybrid system: it does not rely on a single type of artificial intelligence, and it does not leave the decision to a standalone model. Rather, it combines sensors, models trained on data, rules laid down by experts, tools for selecting among alternatives, equipment, and the human being responsible for the decision. This integrated system may consist of the following layers:

  1. Information gathering:
  2. The chain begins with a sensor measuring soil moisture or temperature, or with an image captured by a phone or a drone, or with a record in which the worker notes what happened in the field. This layer does not provide a decision; rather, it conveys a partial description of reality, which may be accurate or incomplete, or affected by a device fault or by the method of measurement.

  3. Checking data quality:
  4. Before the system uses a reading, it checks whether it is plausible. If a moisture sensor reports a sudden jump that does not match rainfall or irrigation, or a device stops transmitting, or some readings arrive in a different unit, those values should not pass to the rest of the system as though they were reliable facts. The system may exclude them, lower its confidence in them, or request that the device be inspected.

  5. Forming an estimate or a forecast:
  6. An analytical model uses the available data to estimate a condition that cannot be measured directly, or to forecast what may happen later. It may estimate the probability of a plant being diseased from an image, forecast changes in soil moisture over the coming hours, or estimate the volume of water demand on the following day. This output remains an estimate, governed by the quality of the data and the limits of what the model has learned.

  7. Applying agricultural and regulatory rules:
  8. Not every predicted result is turned into an automatic recommendation. There may be a rule that prevents irrigation when wind speed exceeds a certain threshold, or prevents the suggestion of a substance not registered for the crop or the region, or requires the case to be referred to a specialist when the diagnosis is inconclusive. These rules are not necessarily learned from data; they are put in place to protect the crop, the worker, and the environment, and to comply with official instructions.

  9. Weighing the available alternatives:
  10. If several possible plans are available, the system uses a computational tool to compare them and choose the most suitable plan within the constraints. It may, for example, try to distribute a limited quantity of water among multiple sectors while taking into account crop needs, pump capacity, operating hours, and energy cost. Here, 'most suitable' does not mean that there is an absolute ideal solution; it means the best solution given the objective and the constraints entered into the system.

  11. Presenting the result, its basis, and its limits:
  12. The interface should make clear to the user what the system is proposing, what data it relied on, how confident it is, and what constraints or missing readings may affect the result. A table that says 'Run irrigation for two hours' is less useful and less safe than a display that specifies the sector in question, the reason for the recommendation, the time of measurement, the status of the sensors, and whether the plan requires approval.

  13. Review or execution:
  14. In some systems, the farmer or supervisor reviews the recommendation before it is carried out. In other systems, commands may be sent automatically to pumps or valves, but within operating and safety limits set in advance. Human beings must retain the authority to stop the system or override its decision when they notice something the data do not capture.

Let us consider the example of irrigation. The process may begin with soil-moisture readings, weather data, and the crop's growth stage. The system first checks the readings to ensure that the sensors are working and that the units of measurement are correct. It then uses a model that tracks changes in moisture over time to forecast each sector's water requirement during the following day. But that forecast alone does not determine the irrigation plan. Agricultural rules may add constraints that prevent adding more water in poorly drained soil, or give priority to a crop passing through a sensitive growth stage. After that, a planning program compares the possible alternatives and allocates pump capacity and operating time across the different sectors. Before execution, the safety layer checks that there is no conflict among the sensors, and that network pressure and valve status permit the plan. It then presents the result to the operator for approval, or carries it out automatically if it falls within the permitted limits. In this example, 'artificial intelligence' did not make the decision on its own. The model provided a forecast, the rules imposed constraints, the planning tool chose among the alternatives, the safety layer reviewed whether execution was feasible, and the operator remained responsible for the final decision. Describing the whole system as an 'AI irrigation model' therefore conceals essential parts of how it works and makes it difficult to determine where the error occurred and who is responsible for it. A model may be accurate, while the system fails at another stage. Examples include the following:

  • A sensor sends a reading in a unit different from the one the system expects.
  • The forecast relies on old data that arrived after the decision window had passed.
  • A result pertaining to the eastern sector is linked to the valve for the western sector.
  • The planning tool chooses a plan that is computationally sound but exceeds the pump's actual capacity.
  • The interface displays a number without stating the time of measurement or the degree of confidence in it.
  • The system sends a correct command, but the valve does not execute it and does not return a signal confirming its status.
  • The decision is executed automatically without giving the operator a clear means of stopping it.

A review of agricultural digitalisation describes an interconnected pathway that begins with sensing, passes through data, context, and decision-making, and ends with action in the real world [SRC035]. The importance of this conception lies in the fact that the quality of each stage depends on what reaches it from the previous stage, and that a small error may be amplified as it moves through the chain. If the moisture reading is wrong, it may produce a wrong forecast. If the forecast is correct but sent too late, the plan may become unsuitable by the time it is executed. If the plan is correct but directed to the wrong valve, a sector that did not need water may be harmed while the sector that did need it remains unirrigated. That is why checking the connections between the system's components may be more important than achieving a slight increase in the accuracy of the algorithm itself. For the in-depth reader: It is not enough to evaluate the model in isolation from the system within which it operates. Error should be traced from the first measurement to the final field impact: how large is the sensor error? How were missing values handled? Did the forecast arrive in time? How was it turned into a decision rule? Did the planning tool take real constraints into account? Was the command executed on the intended equipment? And what agricultural impact resulted from it? A given model may appear better in computational tests, while a system using a simpler model may be more reliable because of the quality of its data, the clarity of its rules, and the strength of its safety layers. Not every layer in the system should become 'intelligent' merely because that is technically possible. In some places, a fixed rule is clearer, easier to test, and safer than a learned model. Preventing a machine from operating when a safety guard is open does not require a model that predicts whether operation is dangerous; it requires an explicit rule that must not be overridden. By contrast, learning may be useful in analysing a complex image or anticipating a change that is difficult to describe with a simple set of rules. Mature engineering does not begin with the question, 'Where can we add artificial intelligence?' Rather, it begins with more precise questions:

  • Which part of the problem needs to learn from data?
  • Which part requires an explicit agricultural or regulatory rule?
  • Which decisions can be carried out automatically, and which require human approval?
  • Where might the connections between components fail?
  • How does the user know that the data are incomplete or that execution did not occur as planned?
  • Who has the authority to stop the system and override it when an unexpected situation arises?

The governing rule: The quality of an agricultural system is not measured by the smartest model within it, but by the safety of the entire chain, from the first reading in the field to the final action carried out there.

11. A Preliminary Guide to Linking the Task, the Method, and Verification

The appropriate method does not begin with the name of the algorithm, but with the question we want to answer and the action that may be based on that answer. The same image may serve as input for classifying the condition of a leaf, locating a weed, or measuring the area of damage, but each of these tasks requires different reference data, produces a different output, and is assessed by different metrics. The following guide is not an automatic recipe for choosing an algorithm, nor does it mean that every agricultural problem belongs in a single category. It is a preliminary map that helps the reader connect six things:

  1. Task: What agricultural question do we want to answer?
  2. Data: What information is needed to answer it?
  3. Method family: What kind of analysis can address the question?
  4. Output: What result should the system provide?
  5. Performance metric: How do we know that the result is useful, rather than merely looking good?
  6. Safety and verification question: What must be checked before the result is turned into a decision or action?

The answer to these questions is not complete until the crop, the environment, the timing of the decision, the cost of error, and the user's ability to respond have been specified. 1\. Classifying the condition of a leaf from an image Task: Determine which class best matches the leaf image, such as healthy, affected by a particular condition, or indeterminate. Required data: Diverse images of healthy tissue and diseased conditions, each linked to an expert-verified reference class or a reliable test. It is not enough for the images merely to be numerous; they must also represent variation in cultivars, growth stages, lighting, imaging devices, and symptom severity. Method family: Image-classification models. Intended output: The name of the most likely condition and a confidence score, with the option for the system to say 'indeterminate' or request an additional image when the evidence is insufficient. What do we measure? An overall success rate is not enough. We should measure the system's ability to detect each important condition, the number of cases it misses, the number of false alerts, and the cases it confuses with one another. Safety and verification question: What does the system do if it encounters a disease it has not learned, symptoms arising from more than one cause, or a poor-quality image? Does it refer unclear cases to a specialist, or force them into one of the known classes? 2\. Locating a weed in preparation for dealing with it Task: Detect the weed in the image and determine its location with enough precision to direct a worker or a machine to it. Required data: Images or video clips showing the locations of weeds and the crop, with reference annotations of their positions. If the system is to guide a machine, it also needs calibration information linking the object's position in the image to its true position in front of the implement. Method family: The work begins with object detection or delineating object boundaries within the image; the result then passes to a control system that determines the implement's movement. Intended output: The weed's location, its boundaries, and the confidence score for its detection, followed by coordinates that the machine can use in time. What do we measure? We measure which weeds were detected and which were missed, the accuracy of their localisation, the speed with which the result arrives, and the number of crop plants damaged during execution. Safety and verification question: Does the implement stop if the system loses confidence, if a person or animal enters the work area, or if it becomes impossible to distinguish the weed from the crop? And has the entire chain been tested, from the camera to the implement's movement, rather than the image model alone? 3\. Yield prediction Task: Estimate the quantity of crop expected to be harvested from a specified area. Required data: Previous yield measurements, and data on weather, soil, cultivar, growth stage, and agricultural practices, with the source of each value and its unit of measurement clearly stated. Method family: Regression models, or models that study how data change over time; spatial models may be added if conditions differ within the field or between regions. Intended output: A yield estimate in an intelligible unit, such as kilograms per hectare, accompanied by a range showing the degree of uncertainty instead of a single number that suggests false precision. What do we measure? Error is measured in the same unit as yield, while also checking whether the stated confidence ranges agree with actual results. If the system says that yield is likely to fall within a certain range, we should test how often that actually happens. Safety and verification question: Has the model been tested on farms, seasons, or regions that were not part of its training? And does the user understand that the prediction is based on past conditions that may not represent an exceptional season or a new cultivar? 4\. Preparing an irrigation schedule Task: Determine when the sectors should be irrigated and how much water each sector should receive, while taking into account the crop's needs and the capacity of the irrigation network. Required data: Soil moisture, weather, crop growth stage, soil properties, water sources, pump capacity, valve status, and time and operational constraints. Method family: The system may combine a model that predicts water requirements, agricultural rules that prevent unsafe decisions, and a planning programme that compares alternatives and allocates the available resources. Intended output: An executable schedule specifying the sector, time, quantity, and operating duration, together with a fallback plan if water becomes scarce or a piece of equipment fails. What do we measure? Success is not measured by reducing water consumption alone, but also by plant condition, yield, energy use, uniformity of distribution, the timing of water delivery, and whether the plan can be carried out with the available equipment. Safety and verification question: What happens when data are interrupted or one of the sensors fails? Does the system switch to manual mode or a safe contingency plan, or does it continue issuing commands on the basis of incomplete information? 5\. Issuing an alert about an animal's condition Task: Detecting a change that may indicate disease, injury, or a disturbance in movement or feeding. Required data: The data may be sounds, images, video clips, or readings of activity, temperature, and feed intake. They should be linked, as far as possible, to cases reviewed by a veterinarian or other specialist. Method family: Detecting unusual cases when we do not have a complete definition of every disease, or classification when we have documented cases from which the system can learn. Intended output: An alert identifying the animal, the time, and the reason for concern, with alerts ranked by priority. What do we measure? The number of important cases the system failed to detect, the number of false alerts per animal or per day, and the interval between the onset of the change and the issuing of the alert. Safety and verification question: Who confirms the condition professionally before treatment? And how many alerts can the team actually examine? A system that sends hundreds of unranked alerts may hide the dangerous case in the noise. 6\. Estimating the area of symptoms and tracking their spread Task: Measuring the affected area in the plant or field, then determining whether the symptoms are expanding or receding over time. Required data: Sequential images or maps captured at different times, together with information about the location, height, angle, and lighting conditions of image capture, so that a fair comparison can be made. Method family: Delineating the boundaries of affected tissues or areas within the image, then analysing how they change across different points in time. Intended output: An estimated area or percentage of damage, the direction and speed of change, together with an indication of the degree of confidence in the comparison. What do we measure? Error in estimating the area, the stability of the result when imaging is repeated, and the system's ability to track the same area from one time to another. Safety and verification question: Did the damage actually increase, or did the light, imaging angle, or camera height change, making the area appear larger or smaller? And are the images being compared on the basis of a single reference standard? 7\. Providing an advisory answer Task: Answering an agricultural question in clear language, drawing on authoritative sources appropriate to the user's context. Required data: The user's question, information about the crop, location, growth stage, and condition, together with reliable and up-to-date sources that the system can consult. Method family: Searching authoritative sources and retrieving the relevant passages, then using a language model to formulate the answer, with rules that prevent claims or recommendations not supported by the evidence. Intended output: An answer that distinguishes between general information and a qualified recommendation, cites its sources, and asks for additional information or declines to answer when the evidence is insufficient. What do we measure? Do the sources support the claims made in the answer? Are the references accurate and verifiable? Is the answer appropriate to the crop, location, and timing, and can the user understand the next step without being pushed towards an unsafe action? Safety and verification question: Does the system refrain from inventing an undocumented dose, substance, official registration, or price? Does it make clear that a product's approval or conditions of use must be verified against the current approved product label and local regulations? 8\. Allocating a limited resource Task: Distributing a limited quantity of water, fertiliser, labour, or machine time among multiple fields or sectors. Required data: The scale of demand, the available resources, priorities, technical and timing constraints, and forecasts relating to yield, cost, or return. Method family: Mathematical planning and trade-offs among alternatives, known in specialised fields as optimisation and operations research. These methods seek a plan that achieves the objective as far as possible without breaching the specified constraints. Intended output: An executable plan, together with a statement of what each party receives and what changes if priorities shift or the resource becomes scarcer. What do we measure? Did the plan respect all constraints? Can it actually be implemented? And what is its effect on cost, yield, time, and fairness among sectors or beneficiaries? Safety and verification question: Who determined the objective that the tool is trying to maximise? And who set the priorities and the minimum acceptable level for each sector? A plan may be mathematically correct yet unfair or agriculturally unacceptable if the objective is reduced to a single number. 9\. Controlling a sequence of actions Task: Choosing a sequence of actions after which the state of the system changes, such as adjusting ventilation and heating in a protected structure, or guiding a moving machine between crop rows. Required data: A description of the current state, the set of possible actions, and data on what happened after each action, together with operating and safety limits. Method family: The system may use conventional control based on rules and known equations, reinforcement learning that learns an action-selection strategy from simulation or experience, or a mixture of the two. Intended output: A strategy that determines the appropriate action for each state, followed by clear operating commands for the equipment. In technical terminology, this strategy is referred to as an 'action-selection policy'. What do we measure? The stability of the system, the speed of its response, the number of times constraints are exceeded, resource consumption, the effects of actions on the plant, animal, or equipment, and its ability to handle unexpected conditions. Safety and verification question: What limits is the system not permitted to exceed, whatever the objective? Who has the authority to stop it? And what safe state does the equipment move to if communication is lost, data fail, or the situation falls outside the range that was tested? This map shows that a single agricultural task may require more than one method, and that success at each stage does not guarantee success for the chain as a whole. Identifying the location of the weed is not enough if the machine cannot reach it safely; predicting water demand is not enough if the irrigation plan exceeds the pump's capacity; and retrieving a correct source is not enough if it pertains to a different crop, country, or date. It also shows that the technical metric and the safety question are not separate matters. The metric indicates how well the system performs in testing, whereas the safety question indicates what may happen when the system makes a mistake in the field. An algorithm may achieve a high average result, while its rare error falls in the highest-risk case. For the reader seeking depth: Choosing the method begins with formulating the decision and the cost of different error patterns, not with comparing the names of models. The appropriate method may change if the intended output changes, even when the data remain the same. The system should also be evaluated at three levels: the quality of estimation, the quality of turning that estimate into a decision, and the quality of implementing the decision in reality. Improvement at the first level does not compensate for failure at the other two. And when evaluating any new application, the matrix can be summarised in the following questions:

  • What specific agricultural question is the system trying to answer?
  • What data does it require, and who verified their accuracy?
  • What output will it produce, and in what unit or with what degree of confidence?
  • What decision or action will be based on this output?
  • Which type of error is more dangerous: a case the system misses or a false alert?
  • Has the system been tested in an independent environment resembling its actual setting of use?
  • What happens when the data are incomplete or the case is unfamiliar?
  • Who reviews the result, and who has the authority to stop it or override it?

The overarching principle: a method is chosen not because it is the newest or the most famous, but because it produces the output the decision requires, can be verified using the appropriate metric, and remains safe when the data are incomplete or the result is wrong.

12. How Do We Compare Models and Systems? From Test Accuracy to Value in the Field

When a developer or company presents a new system, the first question is often: How accurate is it? That is a legitimate question, but it is not enough to judge its usefulness. Two systems may achieve the same score in a laboratory test, yet differ greatly in the kinds of errors they make, their response speed, their operating cost, their ability to function in the field, and the extent of the harm that may occur when they are wrong. Imagine two models, each achieving an overall score of 92% in classifying leaf images. The first may be good at detecting healthy cases, yet miss the early stages of a dangerous disease. The second may detect most diseased cases, but generate so many false alerts that the user loses confidence in it. The overall score appears the same, but the value of the two models on the farm is not the same. The model itself may be good, while the system that contains it remains weak. The model may analyse the image correctly, but the application compresses it until detail is lost, or delays sending the result, or displays it without indicating the confidence level, or links it to a recommendation that is unsuitable for the region. We therefore need to distinguish between three things:

  • Model performance: how good the estimate or classification it produces is.
  • System performance: how well the whole chain performs, from data collection to displaying the result or carrying out the action.
  • Agricultural impact: whether using the system has actually improved the decision and the outcome on the farm.

A comparison is fair and useful only if it addresses the following levels. A. Start with what we do now: the baseline for comparison The baseline is the method we use as a reference point before introducing the new system. Put more simply, we ask not only, 'Does the model work?', but also: Does it work better than what we do today, and by enough to justify the cost of change? The baseline may be:

  • an inspection carried out by a specialist or an agricultural extension officer.
  • a simple rule devised by a local expert.
  • the average yield over previous seasons.
  • the assumption that tomorrow's reading will be close to today's.
  • a fixed irrigation schedule calibrated to local conditions.
  • the older digital system that the new one is intended to replace.

So if the goal is to predict yield, the model can be compared with the field's average yield over previous seasons. If the goal is to detect a disease from an image, it can be compared with the performance of the current inspection, while making clear whether that inspection is carried out by a farmer, a technician, or a specialist. If the goal is to manage irrigation, it can be compared with the schedule currently in use in terms of water, energy, plant condition, and yield. A weak reference method should not be chosen deliberately just to make the new model appear superior. A complex system may outperform a naive rule, yet fail to outperform the farmer's experience or a simple system calibrated locally. In that case, the technical gain is limited and may not justify the cost of hardware, connectivity, maintenance, and training. Nor should the comparison be limited to the question, 'Who achieved the higher number?' It should also include:

  • the amount of improvement over the current method.
  • the consistency of that improvement across fields and seasons.
  • the types of errors that decreased or increased.
  • the cost of achieving the improvement.
  • the extent to which the user is actually able to benefit from it.

The practical meaning of the baseline: first we determine the level of performance we have without the new system, then we test whether the system adds real value beyond it. B. Make the test close to actual conditions of use A model learns from a set of examples called the training data. After learning is complete, it must be tested on independent examples it has not used before. The aim is not to find out how well it remembers earlier examples, but whether it can handle new cases resembling what it will encounter on the farm. If the system will be used across multiple fields, the test should include fields that were not part of training. If it will operate across different seasons, it is not enough to test it on images from the same season. If users take pictures with phones of varying quality, it is not valid to evaluate it entirely on images captured by a single camera under ideal conditions. A random split of the data may give a misleading impression. If ten similar images are taken of the same plant, then some are put into training and the others into testing, the model may recognise the plant, the background, or the imaging conditions rather than the agricultural signs we want it to learn. A stronger test separates, depending on the nature of the task, among:

  • different plants.
  • different plots or fields.
  • different farms or regions.
  • different seasons or years.
  • different varieties and growth stages.
  • different imaging devices or sensor types.
  • different languages or dialects, and different question formulations in conversational systems.

This does not mean that every test must cover the whole world, but that it must faithfully represent the environment in which the system is meant to be used. If the system was tested in a limited environment, those limits must be stated, not its result presented as though it were valid for every crop, farm, and season. Candidate models should also be subjected to the same test, under the same conditions, and with the same data. The comparison is not fair if one is tested on clean images while another is tested on difficult field images, or if one is measured on a fast device and the other on a low-power phone. The practical meaning of independent testing: we put the system in front of new cases that reflect the reality of its use, then see whether it maintains its performance once the conveniences of the laboratory fall away. C. Link the metric to the consequence of error Not all errors are equal. A system may make a mistake by raising an alert about an infection that is not there, or by treating an infected plant as healthy. Both are errors, but their consequences may differ. A false alert may waste inspection time, prompt additional analysis, or annoy the user. Missing a serious infection, by contrast, may delay intervention until the damage spreads. In another application, such as operating a machine, a false alert may be less dangerous than an incorrect movement that harms the crop or the worker. Therefore, the overall average must not conceal the more important type of error. A comparison should ask:

  • How many dangerous cases did the system miss?
  • How many false alerts did it issue?
  • In which varieties or growth stages are errors more frequent?
  • Does it make more mistakes in early cases than in obvious ones?
  • What is the cost of each type of error?
  • Can the team handle the resulting number of alerts?

And if the task is to predict a number, the error should be expressed in a unit that the decision-maker can understand. Instead of saying that 'the average error is 0.12' without explanation, one can say that the yield prediction deviates on average by a certain number of kilogrammes or tonnes per hectare. When planning storage or transport as well, it should be made clear whether the usual small errors conceal a small number of large errors that could disrupt the entire plan. Nor is it enough for the difference between two models to show up in a single number. The advantage may be so slight that it could reverse if we repeated the test in another season or on another sample. Therefore, one should look at the amount of variation across tests, not just the single best result. The practical meaning of the metric: we choose the measurement method according to the decision that will be affected by the result, and translate the error into an effect that the farmer, specialist, or operations manager can understand. D. Test how the system expresses uncertainty, and when it abstains from judgement The model does not possess the same degree of knowledge in every case. The image may be clear and similar to what it learned from, or it may be dark, incomplete, or from a case it has not seen before. Sensor readings may be complete and consistent, or some may be missing while the rest conflict. This variation is called uncertainty: the available evidence does not support all outcomes to the same degree. The system can express this in several ways, including:

  • Displaying a confidence score instead of issuing a categorical judgement.
  • Presenting several ranked possibilities.
  • Displaying an expected range instead of a single number.
  • Indicating missing or conflicting data.
  • Requesting an additional image, reading, or piece of information.
  • Referring the case to a specialist when the evidence is insufficient.

In yield prediction, saying that the expected yield lies between 4.2 and 5.1 tonnes per hectare may be more honest and useful than giving a single figure of 4.7 tonnes, especially if future weather data are unstable. In diagnosing a leaf image, it may be more appropriate for the system to say that the image is insufficient to distinguish between two causes, and then request an image of the underside of the leaf. But the confidence score displayed by the model is not guaranteed truth. It may say it is 80% confident while its actual results are much worse. That is why confidence calibration is examined: testing whether the confidence numbers align with reality. If we gather a large number of cases to which the system assigned confidence of about 80%, its results should be correct at a rate close to that, not correct in only half of them. A case outside the scope of what the system has learned is one that clearly differs from its training examples: a new variety, a different camera, a disease not included in the data, lighting conditions that are sharply different, or a question in a language on which it was not tested. The model may answer confidently even in such cases, which is why its ability to detect that an input is unfamiliar must be tested. Abstention means that the system does not force itself to issue a diagnosis or recommendation when the evidence is insufficient. It is neither a random stoppage nor a sign of failure, but an intentional behaviour, such as saying: 'The current image is not sufficient to determine the condition safely. Take a clearer image of the underside of the leaf, and specify the crop and growth stage, or refer the case to a specialist.' Abstention itself must also be evaluated. If the system abstains in most cases, it may appear safe while being of little use. If it never abstains, it may give confident answers in cases it does not know. What is needed is a balance suited to the seriousness of the task and the team's capacity to review referred cases. The practical meaning of uncertainty and abstention: we ask not only whether the system answered, but whether it recognised the limits of its evidence, whether it asked for help at the right time, and whether that request reached the person able to follow up. E. Measure the system's performance in operation, not only in the demonstration A system may work well in a short public demonstration, or on a powerful computer connected to a stable network, and then run into difficulties when used day after day on the farm. For this reason, the comparison includes aspects that do not appear in an accuracy figure, such as:

  • The time between data entry and the appearance of the result.
  • The system's ability to run on the phone or device that is actually available.
  • Energy consumption and the volume of data transmitted.
  • The need for a continuous internet connection.
  • What happens when the network is weak or the power goes out.
  • The cost of cameras, sensors, servers, and subscriptions.
  • The time required to install, calibrate, and maintain the equipment.
  • How easily the software can be updated without disrupting work.
  • The clarity of the interface and the user's ability to understand the result.
  • The availability of technical support and spare parts.
  • The ability to correct an error or revert to a manual way of working.

A model that performs slightly less well in laboratory testing may be more useful in the field if it runs quickly on an available phone, presents an understandable result, and continues to perform its basic functions when connectivity is weak. Conversely, the more accurate model may be impractical if it requires an expensive device, continuous connectivity, or a response time longer than the intervention window. If a weed-detection system sends the result after the machine has already passed the weed's location, high image accuracy is of no use to it. If a disease alert arrives days after the data are captured, it may lose its value even if the diagnosis is correct. The practical meaning of operational performance: a useful system is one that suits everyday working conditions, not one that succeeds only when ideal demonstration conditions are available. F. Monitor changing conditions and declining performance after deployment A good result when the system is launched is no guarantee that it will continue. Agricultural environments change: new varieties appear, planting dates shift, cameras and sensors are replaced, practices change, and a pest or condition not represented in the original data may spread. When the nature of the data the system receives changes from the data on which it was trained and tested, this is called data drift. The point is not that the data literally move, but that the properties of the reality the system sees are no longer what they were. Examples include:

  • Training a model on images from a specific camera and then replacing it with a camera with different colours and image processing.
  • The cultivated variety changes, so the symptoms appear differently.
  • The climate or the timing of the season changes, so the relationship between weather and yield changes.
  • Modifying the feeding or lighting system in a greenhouse.
  • The method of entering records, or the field names in the application, changes.
  • New questions and dialects appear in the conversational advisory system.

It may also happen that the data remain superficially similar while the relationship between the data and the outcome changes. A particular temperature may have been associated with plant stress under an earlier irrigation system, and then that relationship changed after ventilation was improved or new shading was installed. Therefore, after deployment the system needs monitoring that includes:

  • Comparing new data with the data on which it was tested.
  • Recording unusual cases or cases in which confidence has declined.
  • Reviewing a sample of results in the field.
  • Tracking the number of false alerts and missed cases.
  • Defining a threshold that triggers re-evaluation.
  • Specifying who decides whether to update or retrain the model.
  • Testing the new version before approving it.
  • Retaining the ability to revert to a previous version if a problem appears.

Retraining is not an automatic solution to every decline; errors or undocumented outcomes may enter the new data, and the system may learn from them and become worse. For this reason, updates must go through review, testing, and approval. The practical meaning of data drift: what was valid last season may not automatically remain valid in the following season, so the system needs continuous monitoring, not a certificate of success granted once and for all. G. Measure what happened after the result reached the user Prediction may improve without the agricultural outcome improving. The system may detect disease early, but the alert does not reach the right person. The alert may arrive, but the user does not understand it. The user may understand it, but not have the material, the worker, or the time needed to intervene. The user may carry out the recommendation, but it may not suit the field conditions. For this reason, the entire chain should be tracked:

  1. Was the model's estimate correct?
  2. Did the system turn it into a clear alert or recommendation?
  3. Did the result arrive at the right time?
  4. Did the user understand it and trust it?
  5. Did it change the decision?
  6. Was the decision implemented as planned?
  7. Did implementation lead to a better agricultural outcome?

These levels do not mean the same thing. Irrigation-forecast accuracy may improve, yet water consumption may not fall because the network has a leak. The system may detect an animal condition early, but the team cannot review the alerts until the end of the day. It may suggest a better distribution of labour, but the plan does not take transport between fields into account. When measuring impact, more than one outcome should be considered. The system may reduce water consumption but increase energy use, or raise yield but increase input costs, or reduce manual labour but increase dependence on an external service that is difficult to repair. Not every change in yield or cost should be attributed to the system merely because it occurred after the system was used; the cause may be a better season, a change in prices, or another agricultural practice. If one wishes to establish that the system made a difference, phased implementation, comparison of similar fields or periods, or an appropriate field design that separates the system's effect from other factors may be required. The practical meaning of measuring impact: assessment does not end with whether the answer was correct; it continues by asking whether the answer arrived, whether it changed the decision, and whether the decision led to a better real-world outcome. H. Review fairness, responsibility, and governance The system may achieve a good average, yet work better for one group of users or farms at the expense of another. It may depend on fast connectivity that is unavailable in remote areas, or on a modern phone that not all users own, or on images from large, well-organised farms that do not represent smallholdings, or on a formal language that does not handle local questions and dialects well. Therefore, the results should be disaggregated, where appropriate, by:

  • Farm size and the nature of production.
  • Region and climatic conditions.
  • Variety and growth stage.
  • Type of device or sensor.
  • Quality of connectivity.
  • Language or dialect.
  • The user's experience and ability to use the interface.

Fairness does not mean that the system should give the same result to everyone; rather, it means examining whether a particular group bears more errors, cannot access the service, or lacks a way to contest an incorrect decision. As for governance, it means the rules and responsibilities that govern the system throughout its life cycle, including:

  • Who authorised its use?
  • Who reviews its data and performance?
  • Who bears responsibility for error?
  • Can the user object or request human review?
  • Which decisions may the system propose or carry out?
  • Which decisions must remain in the hands of a specialist or an authorised body?
  • How are records kept and privacy respected?
  • How are changes and versions documented?
  • Who has the authority to stop the system when a risk emerges?

These questions become more important when the system affects plant or animal health, worker safety, the use of water and regulated materials, or the product's acceptance in the market. This perspective is guided by the idea of managing AI risks across the stages of its life cycle, from design and testing to operation, monitoring, and updating [SRC033]. The practical meaning of governance: It is not enough to know what the system can do; we must also know who monitors it, who reviews it, and who has decision-making authority when it makes a mistake. For the advanced reader: A professional comparison does not present a single number as a final judgement. It should make clear the size of the test sample, the degree of variation in the result, performance across different groups and environments, its stability over time, and the cost of each type of error. The conditions of comparison should also be held constant: the same data, the same task definition, the same alert threshold, and the same operating environment. One model may outperform when it is allowed to send many alerts, whereas another may outperform when the number of alerts is limited to what the team can review. Comparison is therefore always tied to the conditions of the decision and the resources available, not to the algorithm's performance in the abstract. A practical checklist for comparing any system The comparison can be summarised in eight questions, but what each question means should be clearly understood:

  1. What decision will change because of the system?
  2. Will the output lead to an additional inspection, irrigation of a sector, treatment of an animal, purchase of an input, or operation of a machine? If we do not know which decision will change, we will not know what performance is required.

  3. What current method is the system being compared against?
  4. Does it outperform the farmer's experience, the specialist's inspection, a simple rule, or the system currently in use? And how much added value does it provide?

  5. Was it tested in an independent environment similar to where it will be used?
  6. Did the test include fields, seasons, devices, varieties, and users whose data were not included in training?

  7. Which error is most dangerous, and how often did it occur?
  8. Is the greatest risk missing the case, issuing a false alarm, delay, or directing the action to the wrong place? And what effect does this error have on the farm?

  9. Does it make the limits of its knowledge clear?
  10. Does it present a confidence level or an expected range? And does it request additional information or refer the case to a human when the evidence is insufficient?

  11. What is the cost of operating and maintaining it?
  12. What devices, connectivity, power, training, support, and maintenance does it require? And what happens in the event of an interruption or breakdown?

  13. How do we detect that its performance has deteriorated?
  14. Who monitors the results after deployment? What sign indicates the need for re-testing or updating? And who approves the new version?

  15. Who has the final decision and the authority to stop it?
  16. Does the user know when the result is advisory? Who reviews high-risk cases? And can automatic execution be stopped and the system returned to a safe mode?

The overarching principle: A model is not better merely because its number is higher; rather, it is better when it outperforms a realistic method, holds up in the environment of use, reveals the limits of its knowledge, operates within the available resources, reduces the most harmful error, and ultimately leads to a better decision and a better outcome on the farm.

13. An Integrated Teaching Case: How Does a Leaf Image Become a Responsible Agricultural Decision?

The following case is an instructional example that brings the chapter's concepts together in a single pathway. It does not provide a diagnosis of any particular disease, nor does it describe a treatment, substance, or dose; rather, it shows how a responsible system should handle incomplete information, and how it moves from an initial image to a decision whose reasons and limits can be examined. A farmer noticed yellow and brown spots on some leaves in one part of the field. He took a picture with his phone, sent it to an assistant system, and then wrote: "What is this problem, and what should I use for it?" The question may seem straightforward, but a responsible answer does not begin with the name of a disease or a substance. A single image may show a symptom—that is, a visible change in the leaf's colour or shape—but by itself it does not establish the cause of the symptom. In the image, the effects of a fungal or bacterial disease, a nutrient deficiency, water stress, chemical injury, sunscald, insect damage, or more than one factor at the same time may look alike. For this reason, the image is the beginning of the examination pathway, not its end. Step One — Turning the General Question into a Question That Can Be Examined The system does not jump from the question "What should I use?" to proposing a treatment. Rather, it restructures the issue into smaller questions:

  • Do clearly describable signs appear in the image?
  • In which part of the leaf are they located?
  • Is the pattern localised or widespread?
  • Is the image sufficient to distinguish among some of the possibilities?
  • What information is missing before moving towards a diagnosis?
  • Is there any sign that calls for prompt review by a specialist?

The system's first response might, for example, be: "The image shows yellow and brown areas, but the image alone is not sufficient to determine the cause or choose a treatment. I will first assess the image quality, and then I need information about the crop, how widely the symptoms are spread, and the surrounding conditions." This opening serves an important function: it separates seeing the symptom from knowing its cause, and prevents the system from turning visual similarity into a certainty it does not possess. Step Two — Verifying That the Image Is Suitable for Analysis Before analysing the spots, the application examines the quality of the image itself. What appears to be a disease symptom may in fact be a shadow, a reflection of light, an artefact of image compression, or a part that is out of focus. The system checks, as far as possible, matters such as:

  • Is the leaf clear or blurred by movement?
  • Is the lighting natural enough to show the true colour?
  • Is the affected area visible in close-up?
  • Is the whole leaf visible, so that the location of the symptoms on it can be determined?
  • Is there an image of the upper surface and another of the lower surface?
  • Is there an image of the whole plant, not just the leaf?
  • Is there a reference object to help estimate the size of the spot?
  • Was the image taken after the leaf became wet or was sprayed with something that might alter its appearance?

By showing the "two surfaces", what is meant is photographing the upper surface of the leaf and its lower surface, because some signs may be clearer on one than on the other. If the image is insufficient, the system should not produce a confident answer merely because the user is waiting for one. Rather, it explains what is missing and asks for the image to be retaken, for example: "The image is unclear around the margins of the spots. Take a closer image in natural light, another image of the lower surface of the leaf, and then add an image of the whole plant and of the affected part of the field." This behaviour is not an obstruction to the user; rather, it protects the user from a result based on an input that does not support a reliable inference. Step Three — Describing What the Vision Tool Sees, Without Claiming a Diagnosis If the image is suitable for analysis, the system may use a computer-vision model to perform one or more tasks:

  • Identifying discoloured areas within the image.
  • Estimating the proportion of the leaf that is affected.
  • Describing the shape of the spots and their apparent distribution.
  • Comparing the pattern with categories the model learned from previous images.
  • Detecting that the image does not sufficiently resemble the cases on which it was tested.

The output must be clear about its limits. So, instead of the system saying: "This is disease X." it can say: "The system detected irregular yellow and brown spots in parts of the leaf tissue. This appearance is visually similar to more than one condition, and the current image is not sufficient to determine the cause." The system may also show the area on which it focused in the image, so that the user or specialist can verify that it analysed the affected leaf tissue, rather than the background or a nearby shadow. As for the expression "unknown category", it does not mean that the model discovered a disease called "unknown". What it means is that the image does not match closely enough the categories that the system learned to distinguish, or that its quality does not permit a safe judgement. In that case, it should display a result such as:

  • The case is not represented among the system's known classes.
  • The image is inadequate for comparison.
  • There are overlapping signs that do not allow classification.
  • A human examination or additional information is required.

Here an important distinction appears: the model may be able to locate the spots accurately without being able to determine their cause. Success in the first visual task does not prove success in diagnosis. Step Four — Adding field context to image context Plant symptoms cannot be separated from the crop, the environment, and management. That is why the system asks for information that could change the interpretation of the image, but it is better not to overwhelm the farmer with dozens of questions all at once. It begins with the most useful questions, then asks for further detail when needed. It may first ask:

  • What are the crop and the variety, if known?
  • What is the plant's age or growth stage?
  • When did the symptoms appear?
  • Are they increasing rapidly, or do they appear stable?
  • Do they appear on a single plant or on a group of plants?
  • Are they concentrated at the edges of the field, in scattered patches, or along the irrigation rows?
  • Did they begin on the older leaves or the newer ones?
  • Are there additional signs on the underside of the leaf?
  • Have symptoms appeared on the stem, fruit, or roots?
  • Has there been heavy irrigation, a water interruption, or a heatwave or cold spell?
  • Has an agricultural input been used, or has the fertilisation programme changed recently?
  • Are there similar symptoms in a neighbouring crop or field?

This information does not add 'incidental details'; it may change the meaning of the image entirely. A distribution of symptoms along an irrigation line may direct the investigation towards water, roots, or irrigation equipment, whereas spread from one focus to neighbouring plants may raise a different set of possibilities. And their appearance on older leaves may carry a different significance from their appearing first on new growth. The system should also distinguish between information entered by the user, information measured by a sensor, and an inference drawn from the image. It should not present them all as though they were facts of equal weight. Step Five — Ranking possibilities instead of declaring a single cause After gathering the image and the context, the system may rank the possible causes. Here, weighing means comparing possibilities in light of the available evidence, not declaring that the highest-ranked cause has been established. It may present the result in this form:

  • One possibility is consistent with the shape and distribution of the spots, but it lacks a sign that should be checked on the underside of the leaf.
  • Another possibility is supported by the recent weather conditions, but it does not fully explain the location of the symptoms.
  • A possibility related to nutrition or irrigation requires reviewing the distribution of the symptoms in the field and soil or water measurements.
  • Other causes are less likely, but they remain possible because the information is incomplete.

Nor is it enough to present bare percentages, because a number may give the reader a greater sense of certainty than the evidence allows. The system should make clear:

  • What evidence raised each possibility?
  • What information lowered it?
  • What data are missing?
  • What next question or test would distinguish between the possibilities most effectively?
  • Have the confidence scores it presents been tested in a similar crop and environment?

In technical phrasing, the term probabilistic layer refers to the part that gathers different pieces of evidence and reorders the possibilities when new information arrives. But this part does not 'know the truth' on its own; rather, it organises what the data support and what they still cannot settle. So if the user adds an image of the underside of the leaf, or a description of how the symptoms are spreading, or the result of a field or laboratory test, the ranking may change. This is not a sign of contradiction, but a natural result of updating the judgement when stronger evidence arrives. Step Six — Searching trusted sources suited to the context After narrowing the possibilities, the system searches approved knowledge sources. By governed sources is meant not just any pages that a search happens to find, but a set of documents whose source, date, scope of validity, and approval status have been specified. These may include, depending on the task and the country:

  • Guidance documents issued by trusted agricultural authorities.
  • Appropriate scientific or diagnostic references for the crop.
  • Updated official bulletins.
  • Instructions issued by the competent regulatory authorities.
  • The current official directions for use for products registered locally.
  • Sampling protocols or referral to a laboratory.

A source does not become appropriate merely because it is correct in itself. It may have been issued by another country where the registered products differ, or it may concern a different variety, or be outdated, or address a disease with similar symptoms in another crop. Therefore, the system should match, as far as possible:

  • The crop and variety.
  • The location, or the country and region.
  • The document date and version.
  • The issuing body.
  • The cultivation type and environment.
  • The crop's growth stage.
  • The regulatory status of the substance or procedure.

The system must present the source in a way that the user or specialist can examine, specifying the document title, the issuing body, the date, and the part on which it relied. It must also distinguish between:

  • What is stated explicitly in the source.
  • What the system inferred by combining several pieces of information.
  • What it did not find sufficient support for.

If it does not find a suitable source, the responsible wording is: "I did not find, within the available set of approved sources, recent and appropriate guidance for this crop and region that would justify a specific recommendation. It is necessary to consult a specialist or the local agricultural authority." Not finding a source is not a gap that the model should fill with guesswork; rather, it is important information that must be shown. Step Seven — Stopping the unsafe transition from probability to treatment After the probabilities and sources have been presented, the safety gate comes into play. By this is meant a set of rules and reviews that prevent the system from turning visual similarity or an unresolved probability into a direct treatment instruction. Different levels of output should be distinguished:

  1. General description: What appears in the image, and what information is missing?
  2. Verification step: What is the next image, examination, or piece of information required?
  3. General low-risk preventive guidance: If it is supported by an appropriate source and does not assume an unconfirmed diagnosis.
  4. Professional diagnosis: This may require a specialist or a reference test.
  5. Selection of a substance, dose, or regulated procedure: This is subject to the diagnosis, local registration, the product's current official directions for use, and the authority of the person authorised to make the decision.

The presence of a disease name in the list of probabilities does not justify moving automatically to the name of a substance. Nor does the presence of a substance mentioned in an old source or one issued by another country establish that it is registered for the crop and region, or that its use is appropriate to the case. The safety gate may determine that the correct output is:

  • A request for additional images or information.
  • A suggestion to inspect the rest of the plant and the field.
  • Referral of the case to an agricultural specialist.
  • Following the specialist's instructions when submitting a sample for examination.
  • Verification with the competent regulatory authority in the country or region.
  • Refraining from proposing a substance or dose because the diagnosis or source is insufficient.
  • Issuing an urgent alert if the available indicators call for rapid professional intervention, while stating the reason for the alert without claiming a definitive diagnosis.

The system should not settle for a general phrase such as "consult an expert" every time, because that makes it of limited use. Better is for it to explain why the matter requires a specialist, what information should be prepared, and what question the specialist needs to resolve. Step Eight — Making the decision, documenting what happened, and following up the outcome After the evidence and its limits have been presented, the decision remains with the farmer or the authorised specialist, depending on the type of action and its level of risk. The outcome may be a request for further examination, monitoring the case for a specified period, conducting a test, or carrying out an intervention approved by the specialist in accordance with local requirements. The system records, with the user's consent and in accordance with privacy rules, items such as:

  • The images and information on which the analysis was based.
  • The time at which the result was issued.
  • The probabilities presented and the confidence levels.
  • The sources retrieved and their versions.
  • The data that were missing.
  • Warnings or reasons for abstaining.
  • The decision taken by the user or specialist.
  • What was actually carried out.
  • Any later results or changes that emerged.

Feedback means the information that returns to the system after the decision: did the specialist confirm a particular cause? Did the test produce a reference result? Did the symptoms change? Was the action actually carried out? But what the user writes afterwards does not automatically become fact on which training can rely. The user may think the diagnosis was correct because they saw improvement, when the improvement may in fact have resulted from a change in weather, irrigation, or some other factor. They may also enter an imprecise name, or attribute the outcome to an action that was not carried out as recorded. Therefore, subsequent information remains of varying strength:

  • The user's observation is useful for follow-up.
  • A specialist's confirmation is stronger than an initial guess.
  • A reference test result may be stronger for a specific task.
  • Improvement after an action does not by itself prove that the diagnosis or the action caused the improvement.

Before this information is used to update the model, its quality must be reviewed, its source verified, whoever approved it identified, and the updated model then tested before release. Otherwise, the system may learn from its old errors and repeat them with greater confidence. What might the answer look like to the farmer? After passing through the initial stages, the system might present an answer such as: "The image shows yellow and brown spots on the leaf tissue, but this appearance can result from more than one cause, and the current image is not sufficient to confirm a diagnosis or choose a treatment. Send a clear image of the lower surface of the leaf, an image of the whole plant, and state the crop and its growth stage, when the symptoms began, and whether they appear on adjacent plants or along the irrigation line. Once this information is complete, the probabilities can be ranked and the next test determined. I will not propose a substance or dose until there is a diagnosis, a source appropriate to the region, and verification against the product's current official directions for use." This answer does not leave the farmer without help, while at the same time not granting them artificial certainty. It clarifies what the system saw, what it could not establish, what information is required, and why it did not proceed directly to treatment. Where can the journey fail? The case shows that the error does not lie in the image model alone. It may occur at any stage:

  • The image may be unusable, but the application accepts it.
  • The model may focus on a shadow or background instead of the spots.
  • The case may not be represented in the training data.
  • The user may enter an incorrect crop name or location.
  • The system may link the image to the context of another field.
  • The system may rank the probabilities with more confidence than the evidence permits.
  • The system may retrieve a correct source, but one that is outdated or not suitable for the country.
  • The interface may summarise the source in a way that changes its meaning.
  • General information may turn into an unauthorised treatment recommendation.
  • The answer may arrive late, after the condition has changed.
  • The system may record a proposed decision as though it had actually been carried out.
  • An undocumented observation may enter the training data as a confirmed diagnosis.

For this reason, it is not enough to say that "the model recognised the image". We need to know how the image became a probability, how that probability was linked to the source, how an unsafe inference was prevented, who took the decision, and how the outcome was documented. For the advanced reader: this case brings together several different functions, each of which should be evaluated separately: image quality, localising the symptoms, classifying the visual pattern, detecting unusual cases, gathering context, ranking probabilities, retrieving sources, verifying their suitability, formulating the answer, applying safety rules, referring to a human, and following up the outcome. The success of one function must not be used as evidence of the success of the others. "Premature diagnostic closure" should also be avoided — that is, stopping at the first explanation that seems convincing — and one should instead seek the evidence that supports it and the evidence that might refute it. If the journey stops after identifying the pattern visible in the image, then we are dealing with an image-analysis tool. If context and the ranking of probabilities are added, then we are dealing with an initial reasoning assistant. If it is linked to appropriate sources, safety rules, a human-review pathway, documentation, and follow-up, then we are dealing with a decision-support system. The difference between these levels does not necessarily appear in the elegance of the interface or the fluency of the answer; rather, it appears in the questions the system asked, the assumptions it disclosed, the sources it made available for scrutiny, and the limits it did not allow itself to cross. The overarching principle: the leaf image may reveal where the question begins, but it does not by itself establish where the answer ends; and a responsible decision is one that connects what the system saw with what it knows about the context, what it can establish from the sources, and what must remain under human judgement.

14. What Should Remain with the Reader from This Chapter?

What should stay with the reader from this chapter? The reader does not need to memorise the names of algorithms, or to know the details of how they are built, in order to think clearly about agricultural artificial intelligence. What matters is seeing the connection between four things: the question we want to answer, the data we have, the result the system produces, and the action that may take place in the field on that basis. A system that analyses an image of a leaf is not performing the same task as a system that predicts yield. A system that forecasts a crop's water requirement is not performing the same task as a planning tool that selects when to switch on the pump. Likewise, issuing an alert is not the same as making a decision, making a decision is not the same as carrying out an action, and carrying out an action alone does not guarantee an improved agricultural outcome. When faced with any application, device, or sales presentation, the reader can organise the assessment into four groups of questions. First — what problem is the system trying to solve? Start with the purpose, not the name of the technology:

  • What specific agricultural problem does the system address?
  • Who is the intended user: the farmer, the extension officer, the veterinarian, the operations manager, or the researcher?
  • Does it describe what is happening now, or predict what may happen later?
  • Does it present probabilities, suggest a decision, or execute an action on a machine?
  • What decision is supposed to change because of the result?
  • Does the output arrive while intervention can still be useful?

A company may say that its system 'detects plant diseases', but that description may conceal very different tasks. Does it identify only the presence of a change in the leaf? Does it compare the image with a limited number of cases? Does it rank possible causes? Does it provide a source? Does it suggest a treatment? Each of these levels involves different data, risks, and responsibilities. One should also ask what the result itself means. If the system displays the number 0.82, what does it mean? Is it a probability? A similarity score? A proportion of the affected area? And if it predicts a yield of 4.7, what is the unit: tonnes per hectare, per feddan, or for the whole farm? And what time period is it referring to? An output whose unit, timing, or relation to the decision is unknown may look precise, but it does not support responsible action. Second — what does the system know, and what can it not know? Every model depends on data it learned from, or on rules and sources it was supplied with. So one should ask:

  • What kind of data was it built on?
  • Where did it come from?
  • Who verified its quality?
  • Does it represent the crops, varieties, seasons, and environments in which it will be used?
  • Does it include early, mixed, and difficult cases, or only clear-cut ones?
  • What information can the system not see or access?
  • What happens if the image is poor, the reading is incomplete, or the sensor is faulty?
  • Does it distinguish between what it actually measured, what the user entered, and what it inferred itself?

A plant-image model may learn from examples, each linked to an expert-verified reference class. A yield model may learn from records linked to reference values measured after harvest. If those reference classes or measurements are inaccurate, the error enters the learning process itself. Nor is the resemblance between the training data and the farm environment a secondary detail. A model trained on images of individual leaves in a laboratory may struggle in a field where leaves overlap and the light changes. A system trained on large farms with regular sensor coverage may not maintain its performance on a smallholding that relies on intermittent measurements. The important question is not only, 'What did the system see?' but also: What can it not see, even though the decision depends on it? An image does not always reveal the condition of the roots, does not by itself explain the history of irrigation and fertilisation, and does not prove the presence of a disease-causing agent. A moisture reading at a single point does not necessarily describe the whole sector. And a digital record does not always show whether the action recorded was actually carried out. Third — how do we know that the system really works? It is not enough for a product to display a high percentage or a successful trial. One should ask:

  • What was the system compared against?
  • Did it outperform the method currently in use, or merely a weak baseline?
  • Was it tested on farms, seasons, or devices that were not part of its training?
  • Which kind of error is more serious in this task?
  • How often did that error occur?
  • Were the results presented for every important case, or hidden inside an overall average?
  • Was only the model shown to succeed, or was the whole system tested?
  • Is there evidence of an effect on the decision and the agricultural outcome, or only on the accuracy of the prediction?

The method currently in use, or the simplest realistic rule against which we compare the new system, is called the baseline for comparison or 'the baseline'. That baseline may be an expert examination, the average of previous seasons, a locally calibrated irrigation schedule, or an older system already in use. The question is not whether the model outperformed nothing, but: Did it add value beyond the realistic method already available, and is that value worth the cost of using it? The test should also resemble real conditions of use. If images of the same plant, or highly similar consecutive images, enter both the training data and the test data, the result may look excellent because the model has effectively seen something like the test before. This is one form of data leakage: information or closely related examples reaching the test set even though they should have remained separate from it. It is also necessary to distinguish between three types of evidence:

  1. Evidence of model performance:
  2. It shows that the model classified images or predicted values well in a specific test.

  3. Evidence of system performance:
  4. It shows that data collection, quality checks, analysis, presentation of the result, communication, and implementation work together under real-world conditions.

  5. Evidence of agricultural impact:
  6. It shows that using the system led to a better decision, reduced harm, conserved a resource, or improved an outcome on the farm.

And the first type does not automatically establish the other two. The prediction may be correct, but the alert arrives late. Or it may arrive on time, but the user lacks the resource needed to carry out the recommendation. And the decision may be carried out, but its effect differs because of a factor the system did not take into account. Fourth — What happens when the system does not know, or when it fails? A responsible system does not behave as though all cases are familiar, all data are complete, and all devices always work. We should therefore ask:

  • Does it make clear that the evidence is insufficient when the case is ambiguous?
  • Does it provide a range or a degree of confidence instead of a definitive number?
  • Can it request additional information?
  • Can it refrain from making a diagnosis or recommendation when it lacks sufficient grounds?
  • What does it do if the sensor, the connection, or the power supply is cut off?
  • Does it switch to a safe mode, or does it continue on the basis of old data?
  • Who reviews unresolved cases?
  • Who has the authority to stop implementation or override it?
  • How does the team detect that performance has deteriorated after purchase or deployment?
  • Who approves updates and new releases?

Uncertainty means that the evidence does not support all outcomes with equal strength. The image may be clear and similar to what the model learned from, or it may be incomplete and depict a case it has not seen before. The sensors may agree, or their readings may conflict. As for abstention, it means that the system does not force itself to issue a judgement when the evidence does not allow it. It may say: "The current data are insufficient to determine the condition safely. Send an additional image, inspect the sensor, or refer the case to a specialist." This is not a sign of weakness if used in the right place. An answer that states its limits and identifies the next step may be more useful than a definitive answer not supported by sufficient evidence. But abstention also requires balance. A system that abstains in most cases may be of little use, and a system that never abstains may be dangerous. What is needed is to specify the cases in which it should answer, the cases in which it should request additional information, and the cases in which authority passes to a qualified human being. When the data or the connection fail, it is not enough for the technical manual to say that the system "supports offline operation". We need to know what actually remains available, when the data were last updated, whether the user can distinguish between a live reading and an old one, and how to return to a safe manual mode of operation. These questions do not oppose innovation The questions above may seem numerous, but they are not meant to obstruct technology or reject it. They help move innovation from an exciting demonstration to a tool that can be trusted, used, and reviewed. A mature technology can explain:

  • The task it performs.
  • The data on which it relies.
  • The conditions under which it was tested.
  • The types of errors it makes.
  • The limits of the result it produces.
  • What happens when data are missing or equipment fails.
  • Who has the final decision.
  • How performance is monitored after use begins.

But a technology that hides these questions behind the name of a famous model, a general percentage, or a dazzling interface still needs maturity, however fluent its presentation. Trust does not arise from the absence of questions, but from the ability of the system and the party responsible for it to answer them with clear evidence. Chapter summary There is no single method that is best for all agricultural problems, nor do AI methods proceed along a ladder that necessarily ends with the largest or newest model. Each method has a kind of question it handles well, a kind of data it needs, and limits that should be understood. What the chapter has presented can be organised according to the function the method performs. Methods that learn from examples with known outcomes In supervised learning, the model learns the relationship between the input and its associated reference outcome. The inputs may be images of leaves, and the reference outcomes may be conditions confirmed by a specialist. Or the inputs may be weather, soil, and management data, and the reference outcomes may be yield values measured after harvest. The strength of this method is that it uses previous examples to predict a new outcome. Its basic limitation is that it learns from the way those examples were constructed, including their errors, biases, and narrowness of representation. As for semi-supervised learning, it draws on a limited number of examples reviewed by a specialist, alongside a larger number of images or records whose outcomes have not yet been determined. This may reduce the review burden, but it does not make up for weak documented examples or the limited diversity among them. Methods that search for patterns without a ready-made answer for every example In unsupervised learning, the system looks for similarities and differences within the data. It may divide the field into zones with similar moisture or yield, or detect records that depart from the usual pattern. But a cluster discovered by the system is not, in itself, an agricultural diagnosis. Two zones may be computationally similar for different reasons. The role of the algorithm is to reveal a pattern that merits examination; interpreting the cause requires field context, inspection, and expertise. Anomaly detection falls within this area in many applications. The system may draw attention to a sensor that has begun to drift, an animal whose activity has changed, or water consumption that has moved outside the usual range. But it does not automatically establish that a disease or malfunction has occurred. Methods that analyse images and change over time Computer vision uses an image or video to classify a condition, determine the location of an object, outline the boundaries of an affected area, or track an object across successive frames. But seeing a spot does not mean knowing its cause, and locating a weed does not mean that the machine can remove it safely. A vision output may be an input to a further stage of inference or control, rather than a complete decision. As for methods that analyse change over time, they examine sequences of readings: how has the moisture changed? Is the animal's activity rising or falling? What is the expected trend in yield or water demand? These methods require attention to seasonality, changing conditions, gaps in measurement, and the need not to assume that a past pattern will always continue. Methods that represent probability and separate prediction from the effect of intervention Probabilistic models help express uncertainty rather than conceal it. They may rank several possible explanations, present a range for expected yield, or update their probabilities when new information arrives. But probability is not useful merely because it appears as a percentage; the reliability of confidence scores must be tested, and it must be shown whether the case resembles the environments in which the system was tested. As for causal inference, it addresses a question harder than prediction. Prediction asks, 'What is expected to happen?' whereas the causal question asks, 'What is expected to change if we intervene?' A model may succeed in predicting yield from soil and management data without establishing that increasing one of the inputs will raise yield. Moving from observing a relationship to recommending an intervention requires stronger evidence, appropriate design, agricultural knowledge, and environmental and regulatory limits. Methods that compare alternatives or deal with sequential actions Mathematical planning and optimisation help allocate resources and weigh trade-offs among plans within clear constraints. They may distribute a limited quantity of water across multiple sectors while taking account of the growth stage, pump capacity, energy, and operating times. But the resulting plan reflects the objective and constraints that the designer specified. If the objective neglects plant health or fairness between sectors, the tool may produce a plan that is computationally successful yet agriculturally unacceptable. As for reinforcement learning, it deals with a sequence of actions that change the state of the environment. The system may choose an action, observe the result, and then adjust its choices to improve a long-term objective. However, this type of learning requires exploring the outcomes of actions, and the system must not try dangerous actions in the field in order to learn that they are unsafe. It therefore requires simulation, gradual testing, explicit safety constraints, and human authority to stop it. Methods that open the door to dialogue with knowledge Language models can summarise texts, translate them, extract information, and formulate an answer in natural language. They make access to knowledge easier, because the farmer can ask in their own language instead of navigating complex menus. But the fluency of an answer does not prove its correctness. A model may produce a persuasive statement without sufficient support. For this reason, it can be linked to a pipeline that searches authoritative sources and retrieves the relevant passages before formulating the answer. Retrieval does not automatically guarantee the correctness of the answer; the system may find a correct source that is nevertheless unsuitable for the crop, the country, or the date. The answer must distinguish between what appears in the source, what the system inferred, and what requires the judgement of a specialist. Tools that help us examine the result Explainable AI tools try to show what influenced the model's result: parts of the image, readings associated with the prediction, similar examples, or a change that might have altered the model output. These tools may help the developer discover an error, or help the specialist review a result. But they explain what the model used, and do not necessarily establish what caused the phenomenon in the field. A persuasive explanation does not turn an incorrect result into a correct one, nor does it make up for weak data or the absence of external testing. The method is part of a larger chain None of these methods operates in a vacuum. The agricultural system usually passes through a chain that begins with reality and ends with a new effect on it: Field reality → measurement or image or record → data-quality check → estimate or prediction → rules and context → decision → implementation → agricultural impact → follow-up The chain can fail at any link:

  • The sensor may produce an incorrect reading or one in an unexpected unit.
  • An image may be linked to an inaccurate reference outcome.
  • Images of the same plant may enter both training and testing, making the result appear better than its true ability to generalise.
  • The model may learn from an environment that does not resemble the farm in which it will operate.
  • The prediction may be correct, but the decision rule may translate its meaning incorrectly.
  • A computational objective may be chosen that overlooks an agricultural or environmental constraint.
  • The system may retrieve a source that does not pertain to the crop or the region.
  • The language answer may appear more reliable than the evidence on which it is based.
  • The alert may arrive after the window for intervention has passed.
  • The correct command may be issued, but the equipment may fail to carry it out.
  • The user may not have the resources or authority needed to apply the recommendation.
  • Reality may change over time while the system continues as though conditions had not changed.

For this reason, comparing two algorithms is not always enough. The difference between them may be small, while the effects of sensor quality, interface clarity, implementation safety, or the team's capacity to respond may be far greater. The contract that links technology to decision The chapter's message can be summarised in the following statement: The technical method is part of a clear contract linking the question, the data, the output, the action, and responsibility. What is meant is not a legal contract, but an intellectual and practical agreement whose terms must be clear:

  1. Question: What specific problem are we trying to solve?
  2. Data: What evidence allows us to answer it, and what are the limits of its quality and representativeness?
  3. Output: What will the system produce, in what units, and with what degree of confidence?
  4. Error: How might it fail, and which error would be most harmful?
  5. Action: What decision or action might be based on that output?
  6. Responsibility: Who reviews, approves, implements, halts, and monitors the impact?

If the algorithm is chosen before these points are understood, it becomes a technical name in search of a problem. But if the design begins with the decision required, the data available, the consequences of error, the realistic basis for comparison, and the operating constraints, then the algorithm can become a precise tool within a larger system that humans lead, review, and hold accountable. A real understanding of agricultural artificial intelligence does not begin when we memorise the names of models, but when we can ask:

  • What did the system actually see?
  • What did it infer?
  • What was it unable to know?
  • What evidence supports the result?
  • What action will it lead to?
  • Who bears responsibility for that action?
  • What happens if the result is wrong?

We benefit from the speed of artificial intelligence when the task is clear, the data are appropriate, and the limits are explicit. We stop to seek a measurement, a source, or a specialist when the result goes beyond what the evidence can establish. Synthesis: Do not ask only, "Is this system intelligent?" Ask instead, "What does it know, how did it come to know it, what does it not know, and what will happen in the field if we trust it?"

Evidence Notes

The two reviews [SRC001] and [SRC002] support the general account of the breadth of artificial intelligence and computer-vision applications in agricultural and food domains, without implying that every application is suitable for all crops and environments. [SRC035] supports viewing the digital agricultural system as an interconnected chain that includes sensing, data, context, decision, and action, which justifies evaluating the system as a whole rather than relying on the model output alone. [SRC006] is used as a bounded example of yield prediction and the use of interpretability tools within a specific research setting. The ranking of factors or the findings from that setting must not be turned into general causal relationships or facts that apply to every field. [SRC009] is used to correct the attribution of weed-detection technology: it concerns the analysis of images from soybean fields using convolutional vision methods, and it does not establish the existence of a robotic removal system based on reinforcement learning. This correction helps distinguish between seeing an object in an image and controlling a machine that interacts with it. The presentation of risks, oversight, and continuous review across the system lifecycle is guided by the framework cited in [SRC033]. As for the integrated teaching case, the guide to linking task to method, the practical questions, and the field examples presented in the chapter, they are an instructional editorial synthesis intended to clarify the distinctions between measurement, analysis, decision, and implementation. These examples do not represent the results of a new experiment, proof of the performance of any particular system, an agricultural diagnosis, or a prescription for treating a crop.

Interactive learning lab

Select AI methods by agricultural purpose

Match the task, data, validation plan, and human oversight before choosing a model.

This enrichment complements the chapter and does not replace its editorial text.

Responsible method selection

Model choice follows the problem and evidence; it does not lead them.

  1. Define the agricultural question
  2. Test data suitability
  3. Evaluate on unseen cases
  4. Apply human oversight

Put this chapter into practice

Choose a situation to see what evidence to check and the responsible next step.

Choose a situation to see what evidence to check and the responsible next step.

Method-task review matrix

Showing 3 of 3 rows.
Method-task review matrix
MethodUseful forCritical review
Supervised learningPredicting a known labelled outcomeLabel quality, drift, subgroup error
Unsupervised learningFinding patterns without fixed labelsStability and agronomic interpretation
Computer visionInterpreting crop or field imageryGround truth, lighting, device and season

Check your method selection

Choose an answer to receive immediate feedback.

Question 1 What should come before choosing an AI model?
Question 2 Why test on unseen cases?
Question 3 Which review is essential for crop-image models?
Score: 0 of 3 correct.

Reader community

Comments and scientific reviews

Contributions are linked to this language and section. Nothing appears publicly until an authorized editor approves it.

Approved contributions

No approved contributions have been published for this chapter yet.

Submit a contribution

Every submission is checked for relevance, safety, and scientific clarity before publication.

Your name, email address, contribution, and book-section context are stored on this site for moderation. Your email address is not displayed publicly, and this book does not retain your IP address or browser identifier with the contribution. Do not include passwords, API keys, phone numbers, or other sensitive personal data.

Only aggregate events are counted. Search terms, comment text, private notes, email addresses, IP addresses, and user-agent strings are never stored in book analytics.