SCILABEL INSIGHTS

Annotating Whole-Slide Pathology Images: Tissue, Tumour, Cell and Biomarker Labels Explained

Introduction

Digital pathology is changing how pathology images are stored, reviewed, analysed, and used for research. Through whole-slide imaging, traditional glass pathology slides can be converted into high-resolution digital images that allow pathologists, researchers, and artificial intelligence teams to examine tissue using computer-based platforms. These whole-slide images (WSIs) can contain millions of pixels and may include complex structures such as normal tissue, tumour regions, individual cells, blood vessels, stroma, necrotic areas, and staining patterns. As artificial intelligence becomes increasingly important in healthcare, these images are also becoming valuable sources of data for developing and evaluating computational pathology systems.

However, simply scanning a pathology slide does not make it ready for AI development. AI systems require structured examples from which they can learn patterns or against which their performance can be evaluated. This is where digital pathology annotation becomes important. Annotation involves identifying and labelling defined regions, structures, cells, or other features within a whole-slide image according to a predetermined project schema. The quality of these annotations can directly influence the usefulness and reliability of the resulting dataset.

What Is a Whole-Slide Image?

A whole-slide image is a digital representation of a physical pathology slide created using a slide scanner. Unlike a conventional microscope, which provides a physical viewing experience, digital pathology allows users to navigate through a scanned slide on a computer. An annotator can begin at a low magnification to understand the overall tissue arrangement and then zoom into specific regions to examine smaller structures in greater detail.

A single whole-slide image can contain many different types of information. Depending on the tissue and staining method, the image may contain normal tissue, tumour tissue, inflammatory cells, necrotic areas, blood vessels, connective tissue, cellular structures, and staining artefacts. This makes annotation a complex task. The annotator must understand what the project requires and distinguish between relevant and irrelevant information. Research into AI-based digital pathology has increasingly demonstrated the potential of whole-slide images while also identifying important challenges involving data quality, validation, bias, and generalisability.

Why Is Whole-Slide Image Annotation Important?

Annotation provides structure to otherwise complex visual information. In a raw pathology image, an AI system does not automatically know which area represents tumour, which area represents normal tissue, or which individual cells are relevant to a particular research question. Human experts can provide this information through carefully designed annotations.

For example, a project may require an annotator to identify the overall tissue area first, followed by tumour regions and then specific cells or biomarker-related features. In another project, only tumour regions may be required. The annotation process therefore depends on the purpose of the dataset and the specific questions that the AI system is expected to answer.

High-quality annotation is particularly important because of the principle known as Garbage In, Garbage Out (GIGO). If a dataset contains incorrect labels, incomplete regions, inconsistent boundaries, or other annotation errors, those problems can affect subsequent AI development or evaluation. Good annotation therefore requires more than simply drawing a shape around something that appears interesting. It requires consistency, attention to detail, clear guidelines, and appropriate quality assurance.

Tissue Annotation

Tissue annotation is generally concerned with identifying areas of the whole-slide image that contain relevant tissue. In a simple project, an annotator may only need to distinguish tissue from the background of the slide. More complex projects may require different tissue categories, such as tumour tissue, normal tissue, stroma, necrotic tissue, or other project-defined regions.

The purpose of tissue annotation is often to establish where meaningful biological information is located. This can help downstream analysis focus on relevant areas rather than empty slide background or scanning artefacts. However, annotators should not automatically create additional tissue categories simply because they can identify them visually. The project’s annotation schema determines which categories are required.

For example, if the project only asks for a general tissue region, an annotator should not independently create additional tumour or cell annotations. Adding information that has not been requested can introduce unnecessary variability into the dataset.

Tumour Annotation

Tumour annotation involves identifying and marking regions that meet the project’s definition of tumour tissue. Depending on the project, this may involve drawing a region around an entire tumour area, distinguishing tumour from non-tumour tissue, identifying specific tumour subregions, or performing more detailed segmentation.

The accuracy of the tumour boundary is particularly important. A rough circle around an abnormal-looking area may not be sufficient. If the guideline requires the annotator to follow the visible transition between tumour and surrounding tissue, the boundary should be drawn accordingly. Including substantial amounts of normal tissue may reduce the quality of the annotation, while excluding genuine tumour areas may result in an incomplete representation of the target.

Tumour annotation also demonstrates why annotators must separate their personal interpretation from the project requirements. A healthcare professional may have a particular clinical opinion about an area, but the annotation should still follow the project’s defined criteria. If the guideline does not resolve an uncertain case, the correct approach is to escalate the issue rather than create an unsupported interpretation.

Cell Annotation

Cell annotation operates at a much smaller scale than general tissue or tumour-region annotation. Instead of marking a large area of tissue, the annotator may be asked to identify individual cells, nuclei, or particular cell types. Depending on the project, examples could include tumour cells, lymphocytes, epithelial cells, macrophages, or other defined cellular categories.

Cell annotation can be challenging because cells may be very small, densely packed, overlapping, partially obscured, or affected by staining variations. Some cells may also have similar visual characteristics, making classification difficult. Consequently, the project guideline needs to clearly define what constitutes a target cell and how ambiguous cases should be handled.

QuPath is particularly relevant to this type of work because it is designed for digital pathology and provides tools for creating annotations and working with objects at different scales. Its documentation describes annotations as regions of interest that can represent large areas as well as smaller structures such as individual cells.

Biomarker Annotation

Biomarker annotation is a more specialised form of pathology annotation. A biomarker is a measurable biological characteristic that can provide information about a disease, biological process, or other biological state. In digital pathology, biomarker-related analysis can involve tissue morphology, cellular characteristics, staining patterns, or other features defined by a research or clinical project.

Artificial intelligence research is increasingly exploring the relationship between digital pathology images and biomarkers. However, annotators must be careful not to make unsupported conclusions. A visually unusual area should not automatically be labelled as a particular biomarker unless the project guideline defines the criteria for doing so.

The annotation instructions should specify what constitutes a positive or negative finding, what visual characteristics should be considered, what region or object should be labelled, and how uncertain cases should be handled. The annotator’s role is to apply these defined criteria consistently rather than create a new interpretation.

Understanding the Difference Between Tissue, Tumour, Cells and Biomarkers

Although tissue, tumour, cell, and biomarker annotations may appear together on the same whole-slide image, they represent different levels of information. Tissue annotation generally identifies the relevant tissue area. Tumour annotation identifies a defined tumour region. Cell annotation focuses on individual cellular structures or specified cell types. Biomarker annotation identifies a defined biological or staining-related feature.

These levels can sometimes be represented as a hierarchy. A whole-slide image may first be divided into tissue and background. Within the tissue, a project may identify tumour and normal regions. Within the tumour, individual cells may then be identified. A further annotation layer may identify biomarker-associated features. However, this hierarchy is not universal. Each project determines which annotation levels are required.

The Importance of Magnification

Successful whole-slide annotation requires appropriate use of magnification. At low magnification, an annotator can understand the overall organisation of the slide and identify where the region of interest is located. At higher magnification, the annotator can examine tissue boundaries, cellular morphology, nuclei, staining patterns, and other smaller features.

A useful workflow is to begin with an overview of the slide, identify the region of interest, zoom into the area, perform the annotation, and then zoom back out to review the annotation in its wider context. This approach helps prevent the annotator from becoming focused on a small area while losing sight of how that area relates to the surrounding tissue.

Using QuPath for Whole-Slide Annotation

QuPath is a widely used open-source platform for digital pathology image analysis. It supports manual annotation and provides tools for working with regions of interest and objects within pathology images. In a basic annotation workflow, the user first opens the correct slide and project, reviews the relevant instructions, examines the slide at an appropriate magnification, identifies the region of interest, selects the approved annotation tool, creates the annotation, applies the correct class, and reviews the result before saving.

The software itself does not determine what should be labelled. That decision comes from the project requirements. QuPath provides the technical environment for creating the annotation, while the project guideline defines what the annotation should represent. This distinction is important because a technically correct annotation in the software can still be scientifically or operationally incorrect if the wrong project rule has been applied.

Common Annotation Errors

One of the most common errors is beginning annotation before reading the project guideline. Annotators may rely on assumptions from previous projects, personal experience, or what appears visually obvious. However, annotation rules can differ significantly between projects, even when the same type of image is being used.

Another common error is including irrelevant tissue within the annotation. For example, an annotator may identify the correct tumour region but create a boundary that includes a substantial amount of surrounding normal tissue. Missing part of the target, creating duplicate annotations, selecting the wrong class, and failing to save the completed work are other avoidable problems.

Annotators should also avoid creating labels that are not included in the project schema. If an object appears relevant but no appropriate label exists, the correct approach is to follow the project’s escalation process. Creating a new category independently can make the dataset inconsistent.

Quality Assurance in Digital Pathology Annotation

Quality assurance is an essential part of digital pathology annotation. A completed annotation should be reviewed before submission to determine whether it meets the project’s requirements. The annotator should confirm that the correct class was used, the correct target was identified, the boundary is appropriate, the annotation is complete, and no unnecessary duplicate or unrelated annotations have been created.

Quality assurance is also about consistency. If two trained annotators are given the same type of case, their annotations should be reasonably consistent when the project guideline provides clear instructions. When substantial differences occur, the project team may need to review the guideline, provide additional examples, or clarify how similar cases should be handled.

Research on AI in digital pathology has highlighted the importance of dataset quality and robust validation. A technically advanced AI model cannot compensate fully for poorly designed or poorly labelled data.

What Should an Annotator Do When Unsure?

Uncertainty is a normal part of complex annotation work. The important issue is how uncertainty is managed. When an annotator encounters a difficult case, the first step should be to review the relevant project guideline and examples. The annotator should then determine whether the available rules provide a clear answer.

If the guideline resolves the case, the annotator should follow it. If the guideline does not provide sufficient information, the case should be escalated according to the project’s review procedure. The annotator should not guess, invent a new label, or create an individual rule.

This approach protects dataset consistency and allows project managers or clinical reviewers to make controlled decisions that can potentially be incorporated into future guideline updates.

From Whole-Slide Images to AI-Ready Data

The journey from a scanned pathology slide to an AI-ready dataset involves several stages. A project may begin with whole-slide image acquisition and preparation, followed by identification of tissue regions. Depending on the project’s objective, annotators may then identify tumour regions, individual cells, cellular structures, or biomarker-related features. These annotations are subsequently reviewed through quality assurance and, where required, independent expert review.

Not every project requires all of these stages. A project focused on tumour segmentation may only require tumour-region annotations, while a project studying cellular morphology may require detailed cell-level annotations. The annotation strategy should therefore be designed around the intended AI use case.

Building Better Healthcare AI Through Better Data

Digital pathology provides an opportunity to transform large volumes of pathology information into structured datasets that can support artificial intelligence research and development. However, the value of these datasets depends heavily on how accurately and consistently the relevant information is identified.

At SCILabel, clinically informed annotation is approached as a structured data-quality process. This means combining appropriate annotation platforms, trained annotators, project-specific guidelines, quality assurance, reviewer feedback, and escalation procedures. The objective is to help transform complex healthcare images into reliable, structured datasets suitable for AI development, evaluation, and research.

Whether the requirement involves tissue segmentation, tumour-region annotation, cell-level labelling, biomarker-related annotation, or broader digital pathology data services, the annotation workflow should always begin with a clear definition of what the dataset needs to represent.

Conclusion

Whole-slide images contain enormous amounts of biological information, but converting that information into useful AI data requires careful human annotation. Tissue, tumour, cell, and biomarker labels each represent different levels of information, and each requires an appropriately defined annotation strategy.

The most important principle for annotators is simple: do not annotate based on assumption. Follow the project guideline, maintain consistent boundaries and labels, review your work carefully, and escalate cases that cannot be resolved confidently.

As computational pathology continues to develop, the demand for reliable, clinically informed, and well-controlled annotation workflows will continue to grow. High-quality annotation is not merely a preliminary step in AI development; it is a fundamental part of building trustworthy healthcare AI datasets.

References

Baxi, V., Edwards, R., Montalto, M., & Saha, S. (2022). Digital pathology and artificial intelligence in translational medicine and clinical practice. Modern Pathology, 35, 23–32.

Mazuco Rodriguez, J. P., Rodriguez, R., Silva, V. W. K., Kitamura, F. C., Corradi, G. C. A., de Marchi, A. C. B., & Rieder, R. (2022). Artificial intelligence as a tool for diagnosis in digital pathology whole slide images: A systematic review. Journal of Pathology Informatics, 13, 100138.

QuPath. (n.d.). How do I draw annotations? QuPath documentation.

Stocken, D. D., et al. (2024). Artificial intelligence in digital pathology: A systematic review and meta-analysis of diagnostic test accuracy. npj Digital Medicine, 7.

Zhang, et al. (2023). Artificial intelligence for digital and computational pathology. Nature Reviews Bioengineering.