Skip to Content

WHAT COUNTS AS GROUND TRUTH IN MICROWAVE BREAST IMAGING

October 8, 2026
Philip Tchatchoua

A machine learning model can only be as certain as the answer it was taught to predict. If an algorithm is trained to detect a tumor, someone must first decide which measurements contain a tumor. If it predicts location, the tumor coordinates must be known. If it estimates size, there must be a reference against which that estimate can be compared.

In machine learning, we usually call this information ground truth. The expression sounds reassuringly absolute. In medical imaging, it rarely is. And in microwave breast imaging, where research can move from numerical simulations to physical phantoms and eventually to patient measurements, the meaning of ground truth changes at every stage.

Ground truth begins easily in simulation

Numerical phantoms offer perhaps the cleanest form of supervision. When a virtual tumor is inserted into a breast model, its properties are known by construction. Researchers may know its exact coordinates, volume, shape, dielectric properties, surrounding tissue, number of lesions, etc. This makes numerical datasets extremely attractive for supervised learning.

A localization model can be trained against exact three-dimensional coordinates. A characterization model can predict the lesion diameter. A reconstruction network can compare its output with the original permittivity map. Recent MRI-derived repositories make these labels more anatomically meaningful by basing the numerical models on segmented clinical examinations rather than purely synthetic geometries.

Yet even here, ground truth is not completely neutral. The tissue segmentation, dielectric property assignment and numerical electromagnetic model all contain assumptions. The location of the virtual lesion may be exact, but the way that lesion interacts with microwaves still depends on the model used to represent it.

Physical phantoms offer a different kind of truth

With physical breast phantoms, the situation becomes more experimental. Researchers manufacture structures representing breast tissues and insert targets at predefined locations. The tumor position and size can therefore be measured directly. This provides strong labels for localization and characterization.

However, the physical experiment introduces uncertainty that did not exist in simulation. A nominally 10 mm inclusion may deform during fabrication. Its dielectric properties may differ slightly from the target values. Temperature can alter material properties. Positioning can vary between scans. The antenna system introduces its own measurement variability.

Recent work on realistic microwave phantoms highlights precisely this need to verify tissue-mimicking materials experimentally rather than assuming that the manufactured object perfectly reproduces its design values. Ground truth is therefore becoming something more useful than an exact number. It becomes a reference value with uncertainty.

One label is not enough for every learning task

The previous article in this series separated microwave breast imaging into three artificial intelligence problems: detection, localization and characterization. Each task requires a different type of ground truth.

  1. Detection: The simplest label may be tumor present / tumor absent. In a phantom experiment, this label is usually unambiguous.
  2. Localization: The model now requires spatial information. Ground truth may take the form of x and y coordinates / x, y and z coordinates / a region of interest / a probability or segmentation map. The definition becomes important. Is the reference point the geometric center of the tumor, its center of mass, or the position of maximum dielectric contrast? These choices can change the measured localization error.
  3. Characterization: Characterization introduces still more complexity. A target can have a nominal diameter, a measured physical diameter, a volume or an irregular three-dimensional shape. The same lesion can therefore have several legitimate descriptions of its size.

For supervised learning, the model needs not only a label, but also a precise definition of what that label represents.

Clinical ground truth is a reference standard

When measurements move from phantoms to patients, the phrase ground truth becomes even more delicate. There is no opportunity to open the breast simply to verify the coordinates predicted by an experimental imaging system. Instead, microwave results must be compared with established clinical information.

Depending on the research question, the reference may come from mammography, digital breast tomosynthesis, ultrasound, magnetic resonance imaging, radiological assessment, histopathology, etc. These sources do not necessarily describe exactly the same thing.

Microwave system output beside the corresponding reference standard information for the same breast. It visually illustrates why clinical validation requires an external reference rather than treating the model prediction itself as evidence. Figure from Álvarez Sánchez-Bayuela et al., BMJ Open, 2024. The article is open access.

Imaging provides an estimate of lesion appearance, location and extent. Pathology provides tissue-level diagnostic information, but only after sampling or surgery. A breast may therefore have several reference measurements, each answering a slightly different question. This distinction becomes particularly important when artificial intelligence moves from detecting an experimental target to predicting clinically meaningful findings.

A current multicentric microwave breast screening protocol provides a useful example of this transition. The study plans to compare microwave-based AI classification with established breast screening information across 10,000 volunteers recruited at nine European centres. The protocol explicitly separates the microwave system output from the clinical reference standard used for evaluation.

Ground truth can contain disagreement

Medical labels are also affected by interpretation. Two experts may disagree on lesion boundaries. Measurements from ultrasound and MRI may produce different lesion dimensions. A biopsy can provide definitive information about sampled tissue, while an imaging examination describes a larger spatial structure.

From a machine learning perspective, these discrepancies matter. If they are ignored, the model is trained as though the labels were perfectly certain. If they are documented, uncertainty becomes part of the learning problem.

This suggests several useful practices: (1) preserve the origin of each label; (2) distinguish measured values from nominal values; (3) record uncertainty when available; (4) avoid combining incompatible reference standards without documenting them; (5) evaluate models separately for clearly defined clinical endpoints.

For quantitative tasks, reporting a prediction error without describing the uncertainty of the reference can give a false impression of precision.

Good validation begins before model training

Ground truth problems are closely connected to dataset design. Suppose several measurements are collected from the same phantom using slightly different antenna positions. If some scans are used for training and others from the same phantom are used for testing, the model may encounter almost identical tissue geometry in both sets.

Performance can then appear stronger than it would on a completely new phantom. The same issue becomes even more important with patient data. Training, validation and test partitions should therefore reflect the level at which generalization is expected.

Possible separation strategies include different measurement scans, different phantom instances, different anatomical models, different acquisition sessions, different patients, different centres, etc.

Recent reviews of AI in microwave medical imaging identify dataset scarcity, heterogeneous acquisition conditions and limited external validation as major challenges for assessing model generalization.

From perfect labels to useful uncertainty

There is a natural progression in the development of microwave imaging AI. In simulation, ground truth can be exact. In physical phantoms, it becomes measurable but uncertain. In clinical studies, it becomes a reference standard assembled from medical observations.

This progression should not be seen as a weakness. It reflects the transition from a controlled engineering problem toward a real medical problem. The important point is to make that transition explicit.

An algorithm trained with perfectly known tumor coordinates should not automatically be expected to behave the same way when clinical lesion boundaries are uncertain. Likewise, performance measured on one phantom configuration cannot by itself demonstrate clinical generalization.

Validation is a ladder rather than a single test

A useful way to think about validation is as a sequence. Numerical models test whether the learning concept works under controlled assumptions. Physical phantoms test whether it survives real electromagnetic measurements. Independent phantom configurations test robustness to changing anatomy and experimental conditions.  Clinical feasibility data test whether the learned signal remains meaningful in patients. Larger prospective studies test whether the final prediction answers a clinically useful question.

Each level asks something different. Passing one does not replace the next. As microwave imaging increasingly incorporates machine learning, the question will therefore shift from How accurate is the model? to a more demanding one: Accurate compared with what, under which conditions, and with how much uncertainty?

That is the question that turns an experimental prediction into credible evidence.

Sources and further reading

  • Álvarez Sánchez-Bayuela, D. et al. Microwave imaging for breast cancer screening protocol for an open multicentric interventional prospective non-randomised clinical investigation to evaluate cancer detection capabilities of MammoWave system on an asymptomatic population across multiple European countries. BMJ Open, 2024. https://doi.org/10.1136/bmjopen-2024-088431
  • Särestöniemi, M. et al. Realistic 3D Phantoms for Validation of Microwave Sensing in Health Monitoring Applications. Sensors, 2024. https://doi.org/10.3390/s24061975
  • Pelicano, A. C. et al. Repository of MRI-derived models of the breast with single and multiple benign and malignant tumors for microwave imaging research. PLOS ONE, 2024. https://doi.org/10.1371/journal.pone.0302974
  • Silva, T. M. M. et al. Machine and deep learning applied to medical microwave imaging from reconstruction to classification. Journal of Physics Photonics, 2025. https://doi.org/10.1088/2516-1091/ae0bd3

About the author

R&D Project Manager | France
Philip Tchatchoua, a graduate in Automation and Industrial Robotics, has strong expertise in Machine Learning, Deep Learning, and project management. With a background in data science, he applies advanced methodologies to solve complex problems and deliver high-quality results.

Leave a Reply

Your email address will not be published. Required fields are marked *

Slide to submit