Akbulut S, Colak C. Multitask learning in hepatocellular carcinoma: Integrating diagnosis, prognosis, and clinical decision support. World J Gastrointest Oncol 2026; 18(9): 121975 [DOI: 10.4251/wjgo.121975]
Corresponding Author of This Article
Sami Akbulut, FACS, MD, PhD, Professor, Surgery and Liver Transplantation, Inonu University Faculty of Medicine, Elazig Yolu 10 Kilometers, Malatya 44280, Türkiye. akbulutsami@gmail.com
Research Domain of This Article
Surgery
Article-Type of This Article
review-article
Open-Access Policy of This Article
This article is an open-access article which was selected by an in-house editor and fully peer-reviewed by external reviewers. It is distributed in accordance with the Creative Commons Attribution Non Commercial (CC BY-NC 4.0) license, which permits others to distribute, remix, adapt, build upon this work non-commercially, and license their derivative works on different terms, provided the original work is properly cited and the use is non-commercial. See: http://creativecommons.org/licenses/by-nc/4.0/
Author contributions: Akbulut S and Colak C conceived and designed the review, performed the literature synthesis, wrote the manuscript, critically revised the content, and approved the final version.
AI contribution statement: A large language model, ChatGPT (OpenAI, San Francisco, CA, United States), was used solely for English-language polishing, including grammar correction, word choice, sentence structure, and readability improvement, as the authors are not native English speakers. No part of the scientific content was generated by AI. The large language model was not used for literature selection, scientific interpretation, study design, data analysis, or formulation of the conclusions. No figures or images were generated by AI. All scientific content was written, carefully reviewed, and approved by the authors, who accept full responsibility for its accuracy and integrity.
Conflict-of-interest statement: All the authors report no relevant conflicts of interest for this article.
Corresponding author: Sami Akbulut, FACS, MD, PhD, Professor, Surgery and Liver Transplantation, Inonu University Faculty of Medicine, Elazig Yolu 10 Kilometers, Malatya 44280, Türkiye. akbulutsami@gmail.com
Received: April 7, 2026 Revised: May 20, 2026 Accepted: June 17, 2026 Published online: September 15, 2026 Processing time: 156 Days and 0.4 Hours
Abstract
Hepatocellular carcinoma (HCC) is a clinically heterogeneous malignancy in which diagnosis, structural characterization, biologic aggressiveness assessment, prognostication, and treatment planning are closely interconnected, yet these objectives are often modeled in isolation. Multitask learning (MTL) jointly optimizes related tasks within a shared representational framework and may improve data efficiency, reduce overfitting, and better reflect the multidimensional nature of clinical decision-making in HCC. This narrative review provides a structured overview of MTL in HCC and current applications. Using a Scale for the Assessment of Narrative Review Articles-informed targeted literature search, 35 studies were examined and 16 were retained for final synthesis, organized into diagnostic, structural, prognostic, and treatment-related tasks. Reviewed evidence indicates that MTL has been applied to tumor segmentation with histological grading, joint prediction of microvascular invasion and vessels encapsulating tumor clusters, recurrence and survival modeling, subtype-specific prognostic stratification, treatment response prediction, and future macrovascular invasion risk assessment. Where directly compared, MTL often outperformed corresponding single-task models and enabled coherent risk stratification. Architecturally, designs have expanded from hard parameter-sharing convolutional neural networks to transformer-based, uncertainty-aware, adversarial, and multimodal models. Although promising for integrated clinical decision support, clinical translation still requires biologically justified task pairing, improved interpretability, robust external validation, prospective multicenter studies, and larger high-quality multitask annotated datasets.
Core Tip: Hepatocellular carcinoma (HCC) management involves interrelated tasks, including diagnosis, tumor characterization, microvascular invasion prediction, recurrence-risk estimation, survival modeling, and treatment-response assessment. Conventional single-task artificial intelligence models address these endpoints separately, limiting their ability to capture shared disease biology. This review highlights multitask learning (MTL) as an emerging integrated framework for HCC analysis. Current evidence suggests that MTL may improve performance across structurally and clinically linked tasks while providing more coherent decision support. We also summarize key architectural trends, current limitations, and future directions required for broader clinical translation of MTL in HCC through prospective validation and multicenter methodological refinement.
Citation: Akbulut S, Colak C. Multitask learning in hepatocellular carcinoma: Integrating diagnosis, prognosis, and clinical decision support. World J Gastrointest Oncol 2026; 18(9): 121975
Primary liver cancer represents a major global health burden, and hepatocellular carcinoma (HCC) is its predominant histologic subtype. HCC is also a prototypical example of a clinically heterogeneous malignancy in which diagnostic, biological, prognostic, and therapeutic questions are deeply interwoven, yet often investigated through fragmented analytical frameworks[1-3]. According to GLOBOCAN 2022, liver cancer is the sixth most commonly diagnosed malignancy and the third leading cause of cancer-related death worldwide, with HCC accounting for 75%-85% of all primary liver cancers[4-7]. Its natural history is often clinically silent, leading to late-stage diagnosis, restricted therapeutic options, and poor prognosis[8-10]. Current international guidelines, such as those from the American Association for the Study of Liver Diseases, the European Association for the Study of the Liver, and the Asian Pacific Association for the Study of the Liver, recommend risk-based surveillance using ultrasonography, with or without alpha-fetoprotein, at 6-month intervals in high-risk populations[11-13]. When a suspicious nodule is detected, multiphasic computed tomography (CT) or magnetic resonance imaging (MRI) enables non-invasive diagnosis of HCC based on characteristic vascular imaging features[9,11-13]. Yet establishing the diagnosis of HCC is only the beginning of clinical decision-making; determining tumor aggressiveness, vascular invasion, recurrence potential, and likely treatment benefit remains substantially more difficult[14-16].
In practice, HCC management is not defined by a single prediction task but by a sequence of linked clinical questions: Whether a lesion is biologically aggressive, whether microvascular invasion (MVI) is likely, whether recurrence risk is high, whether a patient is likely to benefit from a locoregional or surgical strategy, and how prognosis should be estimated before or after treatment[15,17-20]. These judgments inform choices among resection, transplantation, ablation, transarterial therapies, systemic treatment, and surveillance intensity[15,19-22]. They are also informed by heterogeneous yet complementary data sources, including cross-sectional imaging, histopathology, and clinical variables[18,20]. Because these endpoints are biologically related rather than independent, analyzing them separately risks overlooking shared disease mechanisms[18-20]. Conventional radiomics, statistical models, and single-task deep learning (DL) approaches are usually optimized for one endpoint at a time, which constrains their ability to capture cross-task structure and limits efficient use of scarce expert-annotated data[18,23].
Recent advances in artificial intelligence (AI), particularly machine learning and DL, have enabled automated feature extraction from complex medical data and accelerated the development of predictive models in HCC[24-28]. Beyond multitask learning (MTL), prior work has explored a broad spectrum of approaches, including linear models, kernel-based methods, ensemble tree-based methods, convolutional neural networks (CNNs)[29-31], and artificial neural networks, for tasks such as diagnostic prediction, risk estimation, disease staging, treatment response prediction, and survival modeling[24,27,32-34]. Although many of these models perform well within narrowly defined applications, most remain optimized for isolated endpoints rather than clinically linked task sets[3,26,27,35]. More broadly, the conventional single-task paradigm in medical image analysis has often treated related objectives independently, even when intrinsic task correlations are present[35]. This limitation is especially relevant in HCC, where dataset sizes are often modest, annotation is costly, and clinically important targets are strongly coupled across the disease course[26,27,33-35]. In HCC specifically, multitask studies have already combined segmentation with classification, pathological differentiation, or MVI-related prediction, illustrating the clinical linkage among these endpoints[10,35-37].
MTL addresses this limitation by training a single model to perform multiple related tasks simultaneously rather than optimizing each task independently[38]. Instead of learning entirely separate task-specific representations, MTL promotes shared representation learning, allowing the model to exploit overlapping structure across endpoints. Through inductive transfer, this framework can improve data efficiency, reduce overfitting, and enable information learned from data-rich tasks to support related data-scarce tasks. In HCC, such shared representations may capture tumor morphology, vascular features, spatial context, and microenvironment-associated patterns that are relevant across diagnostic, structural, prognostic, and treatment-related tasks, thereby providing a biologically and clinically grounded rationale for MTL. Recent HCC studies provide early support for this premise, showing that multitask DL models can jointly address combinations of tumor segmentation, subtype characterization, MVI prediction, treatment response modeling, and survival estimation[3,7,10,16,30], with improved performance over corresponding single-task approaches in some directly comparative studies[35-37,39-41].
At the same time, adjacent paradigms such as multimodal data integration and federated learning are also becoming increasingly important in HCC research. Multimodal learning seeks to combine imaging, genomics, pathology, and clinical information, whereas federated learning addresses privacy-preserving model development across institutions. These paradigms are highly relevant, but they solve different problems. They improve data fusion or data governance; they do not, by themselves, resolve the problem of jointly modeling multiple clinically interdependent tasks. For this reason, the specific conceptual and translational role of MTL in HCC cannot be inferred from the broader AI literature alone and requires dedicated synthesis[9,29,42-44].
Certain studies cited for contextual or architectural support in this review - particularly those addressing federated learning pipelines, multimodal fusion strategies, or hybrid machine-learning classifiers - do not constitute core MTL evidence in the strict sense. These are included where they illuminate broader architectural trends or provide clinically relevant comparators, and are identified as such throughout the text to preserve conceptual precision[45-47].
Against this background, this review addresses this gap by providing a structured, clinically grounded overview of MTL in HCC. Specifically, we organize the literature according to a clinically oriented framework - MTL tasks in HCC - comprising diagnostic, structural, prognostic, and treatment-related tasks. Within this framework, we categorize current MTL approaches conceptually and methodologically, summarize key architectural developments, evaluate major clinical applications including subtype prediction, MVI prediction, recurrence and survival modeling, and treatment response prediction, and discuss unresolved challenges such as task selection, loss balancing, interpretability, and negative transfer. Taken together, the available literature suggests that MTL is more than a technical refinement of existing AI pipelines and may serve as an integrative framework for personalized HCC decision support, although broader clinical translation still requires prospective investigation (Figure 1).
This article was developed as a narrative review and was structured in accordance with the general principles of the Scale for the Assessment of Narrative Review Articles (SANRA) framework[48]. Accordingly, the methodology was designed not to achieve exhaustive systematic retrieval, but to ensure clear justification of the review’s importance, transparency of the literature identification process, appropriateness of referencing, and a balanced critical synthesis of the selected evidence.
Literature identification strategy
A targeted and concept-driven literature search was performed across Google Scholar, PubMed, IEEE Xplore, and arXiv to identify studies relevant to MTL in HCC. The search focused on publications from 2018 to 2025, reflecting the period in which contemporary DL and multitask modeling studies became more visible in HCC research. Search terms included combinations of “hepatocellular carcinoma” OR “HCC” with “multitask learning”, “multi-task deep learning”, “joint learning”, “simultaneous prediction”, “segmentation”, “classification”, “survival”, “prognosis”, “microvascular invasion”, and “treatment response”. Additional studies were identified through backward citation tracking and contextual review of related publications.
Because the aim of this review was interpretive and clinically integrative rather than exhaustive, study identification followed a relevance-based narrative review approach. Studies were prioritized when they contributed substantially to understanding the conceptual development, methodological diversity, or clinical application of MTL in HCC. Among 35 studies examined during the literature review process, 16 were judged to be most relevant and were therefore included in the final narrative synthesis; these comprised 15 full-text studies and one conference abstract for which sufficient methodological and quantitative performance information was available for limited narrative inclusion.
Eligibility logic and study prioritization
In keeping with SANRA principles, study selection was guided by the relevance of each paper to the central review question, namely how MTL has been used to address clinically meaningful tasks in HCC. Priority was given to studies that examined MTL or closely related joint-learning strategies, particularly when these addressed one or more components of the proposed HCC task taxonomy: Diagnostic, structural, prognostic, and treatment-related tasks. Studies using clinical imaging data (CT or MRI), histopathological or whole-slide imaging, or multimodal clinical data were considered especially relevant when they provided sufficient technical and outcome detail for critical interpretation. Articles were generally not emphasized in the main synthesis when they were not directly relevant to HCC, did not meaningfully engage with multitask or joint-task modeling, or lacked enough methodological detail to support interpretation. Single-task studies were not treated as core evidence unless they were needed to provide clinical or methodological comparison.
To operationalize relevance more transparently, studies were prioritized according to the following criteria. Inclusion criteria comprised: (1) Direct application to HCC using clinical imaging (CT or MRI), histopathological or whole-slide imaging, or multimodal clinical data; (2) Explicit implementation of MTL or a closely related joint-learning strategy (e.g., multi-output optimization within a single model); (3) Clear reporting of at least one quantitative performance metric [area under the curve (AUC), C-index, Dice coefficient, accuracy, or equivalent]; and (4) Sufficient methodological detail to support critical interpretation. Exclusion criteria comprised: (1) Absence of direct relevance to HCC (e.g., studies on other liver diseases or general medical imaging without HCC-specific analysis); (2) No genuine multitask or joint-learning element; (3) Insufficient methodological detail; and (4) Conference abstracts were generally excluded unless they were directly relevant to the review question and provided sufficient methodological and quantitative information for limited narrative inclusion. Studies employing single-task architectures were retained only when they provided essential methodological or clinical comparator context.
Data abstraction and synthesis strategy
For each selected study, the review extracted the elements most relevant to narrative comparison and conceptual synthesis. These included model architecture, task-sharing strategy, loss design, validation setting, study population characteristics, task combinations, clinical purpose, and reported performance metrics such as AUC, accuracy, Dice coefficient, concordance index, and survival-related measures. When available, comparisons with single-task baselines were also considered. The synthesis was conducted as a clinically oriented thematic narrative synthesis rather than a pooled quantitative analysis. Specifically, studies were organized according to the MTL task taxonomy in HCC - diagnostic, structural, prognostic, and treatment-related tasks - and were further interpreted according to whether they contributed primarily as core clinical evidence, architectural innovation, or methodological support. This structure was chosen to preserve coherence, minimize purely descriptive listing of studies, and maintain alignment with the central argumentative aim of the review.
Narrative review quality considerations
To remain consistent with SANRA-oriented narrative review standards, particular attention was given to four elements during writing and synthesis: The justification of the review topic, transparency of literature identification, balanced discussion of the evidence, and critical interpretation rather than simple summary. For this reason, the methodology was intentionally designed to support a focused, analytically structured, and clinically meaningful narrative review rather than to replicate the procedural conventions of a systematic review.
SYNTHESIS OF FINDINGS
The general characteristics of the 16 studies that met the above screening criteria are summarized in Table 1, and their key MTL-related features and areas of application are discussed in detail below[7,10,16,18,30,36,37,40,49-56]. A comprehensive evidence summary organized by primary task category, imaging modality, approximate sample size, multi-center design, external validation status, and key study-level limitations is additionally provided in Table 2[7,10,16,18,30,36,37,40,49-56].
Table 1 Overview of multitask learning models in hepatocellular carcinoma.
MVI AUCs were 0.918 training, 0.800 internal, and 0.837/0.815/0.800 external; RFS C-index was 0.763 training, 0.716 internal, and 0.628/0.675/0.728 external; PA-TACE benefit was observed only in the predicted high-MVI-risk/Low-survival-score subgroup
Detection ACC 93.33%, SEN 93.15%, SPE 93.71%, IoU 82.93%; size-grading ACC 77.78%-96.87%; center-point MAE 2.74 mm; max-diameter MAE 3.17 mm; area MAE 144.51 mm2
Multiphase Gd-EOB-DTPA-enhanced MRI: Late arterial, portal venous, hepatobiliary phases
MTL improved MVI prediction from AUC 0.896 to 0.917 with ACC 90.0%; VETC was predicted with AUC = 0.8604 and ACC 82.5%; MVI/VETC-based stratification was associated with OS/RFS
Future macrovascular invasion prediction + OS stratification
MTnet with segmentation subnet + clinical/radiological/radiomic fusion
Portal venous phase CT
Combined model CR-DR achieved AUC 0.877 in training and 0.836 in external validation; model-defined risk groups stratified time to macrovascular invasion and OS
Contrast-enhanced CT: Late arterial + portal venous phases
Test AUCs were 0.79 for MVI and 0.72 for Edmondson-Steiner grade; OS C-index was 0.73, with 3 years, 5 years, and 10 years AUCs of 0.85, 0.90, and 0.89
Retrospective multicenter study; 10-fold CV in center 1 + external independent test set from centers 2 and 3
Yes
Limited sample size; single-center training with potential inter-center domain bias; HBP-only model; no clinical-variable integration; no prospective validation
Single-center retrospective design; no external or prospective validation; arterial/portal-venous MRI only; clinical/genomic variables not integrated; sequential pipeline without demonstrated joint multitask optimization or shared-representation learning
Retrospective multicenter study; training/internal test split + three external test sets
Yes
Retrospective design; predominantly HBV-related Chinese cohort; moderate RFS performance in some external sets; nonrandomized PA-TACE benefit analysis; no prospective validation
Retrospective single-center study; 10-fold CV for ER prediction + 5-fold CV for FLL classification
No
Small single-center HCC cohort; no external or prospective validation; ROI-based 2D slice analysis; multitask component limited to pre-training; method restricted to multi-phase CT images
Retrospective multi-institutional study; training/internal test cohorts from one institution + external test cohort from four centers
Yes
Retrospective design; MTM cohort mainly surgical; HAIC-only prognostic cohort; limited generalizability to other treatment contexts; no prospective validation; whole-liver ROI approach not compared with tumor-specific ROI; complications during and after HAIC or TKI treatment were not analyzed
Retrospective single-center study; surgical cohort for histologic-score development + TACE cohort for survival modeling
No
Survival model limited to TACE-treated patients; retrospective design; no external validation; complex multi-stage pipeline introduces risk of compounding overfitting; label reliability for histological surrogates not formally assessed
The rationale behind MTL rests on the principle of inductive transfer: Knowledge acquired while solving one task can facilitate learning of a related task. When tasks share underlying visual or semantic features - such as tumor location, size, shape, and tissue context - a jointly trained model may learn more transferable representations than one optimized for a single objective. This principle is particularly relevant when multiple clinical endpoints depend on overlapping image-derived features rather than independent information streams. In practice, MTL architectures fall into two broad categories. In hard parameter sharing, all tasks share a common set of hidden layers (the encoder), while each task maintains its own output head. This design promotes common feature extraction while preserving endpoint-specific predictions. In soft parameter sharing, each task has its own network, but the models are encouraged to learn similar representations through cross-network regularization or shared latent spaces. The choice between these strategies depends on the degree of task relatedness and the specific demands of the clinical application. Accordingly, hard parameter sharing is often preferred when tasks are strongly related and data are limited, whereas soft parameter sharing may be advantageous when task interaction must be preserved without forcing all tasks into a single shared representation. In HCC, most currently available studies rely on hard parameter sharing, whereas more recent designs increasingly introduce selective, expert-guided, or uncertainty-aware sharing strategies to better manage task interference and improve generalizability[35,51].
MTL offers several potential advantages for medical image analysis. By learning from multiple related tasks simultaneously, the model is exposed to richer and more diverse supervisory signals, which may reduce the risk of overfitting - a critical concern when labeled medical datasets are small. Knowledge transfer between tasks allows the model to leverage larger-sample or more densely annotated tasks to support smaller-sample endpoints. This is clinically relevant in HCC because segmentation, lesion localization, vascular invasion prediction, recurrence modeling, and treatment-response assessment may share tumor-level and peritumoral information, even though they represent different clinical questions. The multi-objective training setup can also act as an implicit regularizer, discouraging the network from memorizing task-specific noise and encouraging extraction of more broadly useful features. A further practical advantage is that a single well-validated multitask model may be more efficient than maintaining separate models for each endpoint, provided that task-specific performance is not compromised. These advantages are particularly relevant in HCC, where cohort sizes are often modest, annotation is costly, and clinically meaningful endpoints are biologically interdependent[35,41].
Despite these benefits, implementing MTL in medical imaging is far from straightforward. The most fundamental requirement is that the tasks being learned together should be genuinely related; when tasks conflict or are only weakly correlated, joint training can produce negative transfer, in which the auxiliary task degrades performance on the primary objective. Even among related tasks, differences in data complexity, label distribution, and convergence speed can create optimization imbalances. Finding appropriate loss weights - so that no single task dominates the shared representation - remains an active area of research. Network architecture also requires careful design: The model must be flexible enough to capture task-specific patterns while still enabling efficient feature sharing. In addition, the inherent noise and uncertainty in medical images (motion artifacts, low contrast) can further complicate multitask optimization. For HCC applications, these challenges are particularly important because structurally dense tasks, such as segmentation, are often combined with clinically sparse endpoints, such as recurrence, survival, or treatment response. Accordingly, the success of HCC-oriented MTL should be judged not by the number of outputs a model produces, but by whether biologically coherent tasks can be jointly optimized while preserving efficient knowledge transfer, task-specific sensitivity, optimization balance, and transparent performance reporting. Data uncertainty and medical image noise may further impair MTL performance, particularly when small cohorts, heterogeneous acquisition protocols, and imbalanced labels coexist[35,39].
Clinical applications of MTL in HCC according to the four-task framework
Diagnostic tasks: In HCC, MTL has been applied to diagnostic and subtype-oriented problems, particularly identification of the macrotrabecular-massive (MTM) variant, which carries a notably poor prognosis. He et al[18] developed a multitask DL radiomics model that simultaneously predicted the MTM subtype and overall survival in patients receiving hepatic arterial infusion chemotherapy (HAIC). Built on a three-dimensional (3D) MobileNetV1 backbone, the model extracted quantitative radiomics features from preoperative CT images and combined them with clinical variables to identify aggressive subtypes before treatment, potentially informing risk stratification rather than directly establishing treatment decisions. This study is noteworthy because it connects subtype recognition with prognostic stratification rather than treating them as entirely separate analytical objectives.
Structural tasks: A major structural application area is the non-invasive prediction of MVI and vessels encapsulating tumor clusters (VETC), both of which are histopathological features with strong prognostic significance. MVI - the presence of tumor cells in small hepatic vessels - is an independent predictor of early recurrence after curative treatment, while VETC is associated with unfavorable outcomes[49,55,57,58]. Chu et al[49] developed a 3D CNN-based multitask model that simultaneously predicted MVI and VETC status from preoperative gadolinium ethoxybenzyl diethylenetriamine pentaacetic acid-enhanced MRI data. The multitask design improved MVI prediction accuracy over single-task baselines by leveraging complementary information from the VETC prediction branch. More recently, He et al[58] developed a transformer-based DL framework integrating radiomic and clinical features for direct three-class MVI classification (M0, M1, M2) from preoperative gadobenate dimeglumine-enhanced MRI, achieving a macro-average AUC of 0.886 in external validation of 132 patients. Furthermore, Zhang et al[59] proposed a DL diagnostic model for MVI on whole-slide histopathology images that provided spatial information essential for accurately predicting HCC recurrence after surgery. In a related extension of biologic aggressiveness modeling, Zhao et al[55] developed a multitask framework based on expert sharing, spatial transformation, and relational reasoning to jointly predict MVI and cytokeratin 19 expression positivity from gadolinium ethoxybenzyl diethylenetriamine pentaacetic acid-enhanced MRI.
Prognostic tasks: MTL has also shown early value for prognostic modeling in HCC. Song et al[16] investigated early postoperative recurrence prediction through multitask self-supervised pre-training, combining two pretext tasks - phase shuffle (capturing intra-image features) and case discrimination (capturing inter-image relationships) - and demonstrated superior performance compared with conventional transfer learning from natural images. Wang et al[51] developed a transformer-based MTL model that jointly predicted MVI and recurrence-free survival (RFS) from preoperative MRI scans of 725 HCC patients across seven institutions. In addition, Liu et al[50] developed a hybrid multi-modal, MTL framework that simultaneously predicted MVI status, Edmondson-Steiner histological grade, and survival in HCC patients undergoing transarterial chemoembolization (TACE). Similarly, Wang et al[52] employed a neural multitask logistic regression model trained on 2197 HCC patients from the Surveillance, Epidemiology, and End Results database, achieving AUC values of 0.824 for 1 year, 0.813 for 3 years, and 0.803 for 5 years survival predictions. Together, these studies indicate that MTL in HCC is not confined to image-based joint classification but also extends to recurrence modeling and survival-oriented integrated frameworks. Collectively, these studies suggest that MTL may generate clinically relevant prognostic information for individualized treatment planning and surveillance in HCC, although prospective validation remains necessary before clinical deployment can be justified[10,16,18,35,51].
Treatment-related tasks: Treatment-related applications of MTL are also emerging. He et al[18] linked MTM subtype prediction with prognosis in patients receiving HAIC, thereby connecting pre-treatment characterization with treatment-relevant stratification. Li et al[10] combined TACE treatment-response prediction with tumor segmentation in a unified MTL framework, facilitating patient-level risk stratification. Fu et al[7] proposed a multitask deep neural network for predicting future macrovascular invasion, a finding that is critical for staging and treatment selection. These studies are clinically meaningful because they move MTL closer to preoperative and treatment-oriented decision support, where invasive risk, response prediction, and recurrence estimation may influence resection strategy, surveillance planning, and adjuvant treatment consideration.
Cross-cutting methodological and translational issues in MTL for HCC
DL architectures in MTL for HCC: CNNs form the backbone of many MTL systems developed for HCC-oriented prediction and segmentation tasks. Their ability to learn hierarchical spatial features makes them well suited for joint optimization across imaging endpoints. For volumetric CT and MRI data, 3D CNNs are particularly useful because they can preserve cross-slice spatial information relevant to tumor segmentation and invasive feature prediction, including MVI and VETC. Lightweight architectures such as MobileNetV1 have also been adopted in MTL frameworks to balance predictive performance with computational efficiency. Transfer learning can further support CNN-based models when HCC imaging cohorts are small, although its benefit depends on domain similarity and validation design. Within the reviewed HCC literature, CNN-based shared encoders remain an important architectural starting point because they can support both structural tasks and endpoint-oriented prediction within the same framework without requiring completely separate models for each task[49,60].
More recently, vision transformers have been introduced into MTL frameworks for HCC, either as standalone encoders or in combination with CNNs. Transformers excel at capturing long-range dependencies through self-attention, making them well suited for tasks that require global contextual information - for example, assessing overall tumor morphology relative to the surrounding liver parenchyma. Wang et al[51] demonstrated the value of a transformer-based MTL model for joint MVI prediction and survival analysis from preoperative MRI. Hybrid CNN-transformer architectures represent a promising but still early direction. In such models, the CNN branch extracts local features, while the transformer branch captures global spatial relationships, and the two streams are fused before task-specific heads. This architectural trend is important because HCC-related targets often depend on both fine local cues and broader contextual information. Li et al[10] designed a multi-component DL network with separate encoder, prediction, and segmentation modules; this architecture achieved strong performance on both treatment-response prediction and tumor segmentation after TACE. Outside the strict MTL paradigm, other hybrid machine learning approaches have also been explored for HCC classification. For instance, Ali et al[61] proposed a linear discriminant analysis-genetic algorithm-support vector machine model that combines linear discriminant analysis for dimensionality reduction with a genetically optimized support vector machine, achieving competitive diagnostic accuracy with reduced computational cost. Although this model does not employ MTL per se, it illustrates the broader trend toward integrating complementary algorithmic components for improved HCC classification. Even when not formally multitask, such hybrid approaches provide useful architectural context for understanding why feature fusion and modular design have become increasingly prominent in HCC AI research[31,61,62].
An emerging direction in MTL architectures for HCC involves integrated multitask frameworks that combine segmentation with histological grading using fused multi-phase MRI data. Such frameworks leverage DL-based segmentation as a feature extraction backbone, feeding spatially informed representations into radiomics-based grading modules, thereby achieving synergistic performance gains across both tasks[30]. Additionally, the integration of attention mechanisms with multi-phase imaging may support visual interpretability; for example, attention heatmaps can reveal that tumor margins and peritumoral areas are the most salient regions for MVI prediction, offering radiologists visual explanations that may improve trust in model outputs when validated appropriately[58]. This observation is particularly relevant in HCC because peri-tumoral context and tumor-border features are often closely linked to biologic aggressiveness. An emerging class of architectures explicitly models cross-task consistency through adversarial mechanisms. One representative example - the task relevance driven adversarial learning framework - is discussed in detail in the Technological Innovations section below[54].
Advantages of MTL in HCC: (1) Simultaneous task learning: A defining strength of MTL is its ability to optimize multiple related objectives within a single model, provided that these objectives share meaningful biologic, anatomic, or clinical information. For example, Wen et al[36] demonstrated that jointly training segmentation and pathological grading tasks yielded better results in both objectives compared with training each task independently, thereby improving overall diagnostic accuracy for HCC. This is particularly important in HCC, where structural, biologic, and prognostic signals are often tightly coupled rather than independent; (2) Improved task-specific prediction through auxiliary supervision: Incorporating segmentation as an auxiliary task provides the classifier with spatially grounded feature maps that highlight tumor boundaries, leading to more discriminative representations. Li et al[10] showed that when expert-guided segmentation was used as a companion task, both tumor grading accuracy and MVI prediction improved compared with classification-only baselines. Thus, segmentation in MTL should not be viewed only as an additional output, but also as a mechanism for lesion-aware representation learning; (3) Greater data efficiency and potential generalizability: MTL can also improve performance in settings where annotated datasets are limited. Because several HCC studies rely on retrospective cohorts with modest sample sizes, shared learning across related tasks can act as an implicit regularizer and reduce overfitting. This advantage is highly relevant in HCC, where manual annotation is expensive and clinically informative labels such as MVI, recurrence, or treatment response are often sparsely available. However, this benefit should be interpreted as conditional, because generalizability still depends on cohort size, label quality, imaging protocol consistency, and external validation; and (4) Potential clinical workflow integration: From a translational perspective, MTL offers the possibility of integrating multiple clinically relevant outputs within a single model. Instead of deploying separate tools for lesion delineation, aggressiveness estimation, recurrence prediction, and treatment-response assessment, one multitasks pipeline may provide these outputs in a more efficient and clinically coherent way. This potential is attractive in HCC, but it remains investigational until validated in prospective, multicenter clinical workflows.
Addressing uncertainty in MTL: Handling uncertainty: Noise, artifacts, and inter-scanner variability in medical images can propagate through shared representations and degrade multitask performance. Xie et al[37] proposed a triplet-uncertainty framework that explicitly models three sources of uncertainty - aleatoric (data noise), epistemic (model uncertainty), and task-level uncertainty - within a multitask network. By weighting each task’s loss according to its estimated uncertainty, the framework achieved improved robustness for both HCC segmentation and malignancy classification on noisy imaging data. This study is conceptually important because it shows that successful MTL depends not only on task pairing, but also on the model’s ability to adapt to differences in task reliability during training.
Technological Innovations in MTL for HCC: (1) Feature fusion and attention mechanisms: State-of-the-art MTL architectures for HCC increasingly employ multi-scale feature fusion and boundary-aware attention modules. Multi-scale fusion aggregates related features at different spatial resolutions, allowing the model to capture both fine-grained boundary details and coarse-level contextual information. Boundary-aware attention explicitly focuses the network on tumor margins, reducing over-segmentation and improving delineation accuracy - benefits that propagate to downstream classification tasks through the shared encoder[36]. These mechanisms are particularly valuable in HCC because peri-tumoral morphology and tumor-border features are often closely linked to biologic behavior and aggressiveness; (2) Ensemble and modular learning techniques: Several groups have integrated modular attention mechanisms - such as selective kernel modules and squeeze-and-excitation blocks - into MTL encoders. These modules allow the network to adaptively recalibrate channel-wise and spatial feature responses, enabling more effective feature aggregation across scales. Wang et al[53] demonstrated that a hybrid network incorporating such modules achieved state-of-the-art results for HCC segmentation in hematoxylin and eosin-stained whole-slide images. Although this work is pathology-oriented, it illustrates a broader architectural trend in HCC MTL toward more selective, scale-aware, and modular feature learning; and (3) Task-interaction-aware learning: A notable architectural innovation is the task relevance driven adversarial learning framework, which employs a modality-aware Transformer for multi-modality MRI feature fusion combined with a task-relevance-driven discriminator for adversarial learning. This framework simultaneously performs HCC detection, size grading, and multi-index quantification, achieving 93.33% detection accuracy on 135 subjects[54]. The adversarial training strategy enforces higher-order consistency among multitask labels, providing a principled approach to mutual task promotion that goes beyond conventional loss-weighting methods. This is an important development because it reflects a shift from simple shared-backbone MTL toward more explicit modeling of cross-task consistency and task interaction.
Challenges and future directions in MTL for HCC
Despite its promise, MTL for HCC prediction, segmentation, and prognostic modeling faces several practical barriers. First, selecting appropriate auxiliary tasks requires genuine biologic and clinical relatedness. The auxiliary objective must be sufficiently connected to the primary task so that shared representations capture meaningful disease structure rather than irrelevant variation. Otherwise, negative transfer may occur, with the auxiliary task degrading rather than improving primary-task performance. This issue is particularly important in HCC, where some task pairs, such as segmentation plus pathological differentiation or MVI plus recurrence risk, are intuitively coherent, whereas others may be much less compatible.
Second, balancing optimization across tasks remains difficult. Learning rates, loss magnitudes, label distributions, and convergence speeds may differ substantially across tasks, making it difficult to ensure that all objectives are learned appropriately. Gradient-based balancing methods and dynamic loss-weighting strategies (uncertainty-weighted losses, GradNorm) have been proposed, but no universally optimal solution exists. The current HCC literature already reflects this challenge, with different studies relying on dynamic weighting, uncertainty-aware loss modeling, or task-specific design choices to prevent one objective from dominating the shared representation.
Third, generalizability remains a major translational barrier. Marked inter-patient and intra-tumor heterogeneity in HCC, together with differences in scanner type, imaging protocol, treatment strategy, and population characteristics, make it difficult to build MTL models that generalize robustly across institutions. This problem is compounded by the scarcity of large, high-quality, multitask-annotated datasets. Structural labels, pathology labels, treatment outcomes, and long-term follow-up are rarely all available in the same cohort, which limits the development of clinically comprehensive MTL systems[37,39,63].
A closer examination of the individual studies reviewed here reveals study-specific limitations that merit explicit acknowledgment. He et al[18] used retrospective multi-institutional data with external testing, but the MTM component was mainly surgical, the HAIC prognostic cohort was treatment-specific, and prospective validation was not available. Chu et al[49] demonstrated robust MVI and VETC prediction, but the study was single-center, included a limited sample size, and used strict exclusion criteria. Wang et al[51] achieved multicenter validation across seven institutions, yet the external validation C-indices for recurrence-free survival remained moderate (0.628-0.728), reflecting the inherent difficulty of generalizing survival models. Song et al[16] demonstrated advantages of self-supervised pre-training but used a single-center HCC cohort and region of interest -based two-dimensional; slice analysis, limiting statistical power and clinical generalizability. Li et al[10] combined treatment-response prediction with segmentation in a retrospective two-center design with external testing, but the small drug-eluting bead-TACE subgroup and the reported Dice coefficient (73.6%) indicate room for further validation and segmentation improvement. Liu et al[50] achieved meaningful survival prediction (C-index 0.733) but relied entirely on a TACE-treated population, restricting generalizability to other treatment contexts. Fu et al[7] reported an external validation AUC of 0.836, but the study remained retrospective and population-specific, limiting inference beyond the original clinical setting. Wang et al[52], drawing on the Surveillance, Epidemiology, and End Results registry, achieved strong population-level predictions but lacked imaging data, limiting mechanistic interpretability. Across these studies, common methodological limitations include retrospective study design, single-center recruitment in several cohorts, modest sample sizes relative to model complexity, incomplete or absent independent external validation in some studies, and potential label noise from retrospective annotation of histopathological endpoints. Collectively, these limitations underscore that current MTL evidence in HCC - while promising - should be interpreted with appropriate caution, and prospective multicenter validation is necessary before clinical deployment.
Fourth, interpretability is essential for clinical adoption but remains insufficiently addressed. MTL-DL models are often treated as black boxes, which limits clinician trust, particularly in high-stakes settings such as cancer diagnosis, recurrence-risk prediction, and treatment selection. Although some studies have begun to use attention heatmaps, feature visualization, or pathology-based spatial quantification, explainability remains less mature than would be desirable for routine clinical deployment. Without clearer mechanisms for explaining why a model predicts MVI, recurrence, or treatment response, widespread integration into HCC care will remain difficult.
Several research directions are likely to shape the next generation of MTL models for HCC. One of the most important is self-supervised pre-training, which learns visual representations from unlabeled data before fine-tuning on downstream tasks. This approach has already shown benefits for HCC recurrence prediction and focal liver lesion classification, outperforming conventional ImageNet-based transfer learning. Uncertainty-aware training - exemplified by the triplet-uncertainty framework - provides a principled way to handle noisy imaging data by dynamically adjusting task weights according to estimated prediction confidence, improving both tumor segmentation and MVI-related prediction. Dynamic loss-weighting strategies, such as the dynamic weighted average, automatically balance competing objectives during training and reduce the need for manual hyperparameter tuning[16,36,37]. These approaches suggest that the future of HCC MTL will depend as much on improved training strategies as on novel architectures. Hybrid DL approaches that combine transfer learning with task-specific fine-tuning have shown encouraging results for histopathological image classification and could be extended to MTL settings. Furthermore, scale-aware multi-instance learning methods offer a pathway toward classifying pre-cancerous lesions from non-invasive imaging, potentially enabling earlier identification of patients at risk for HCC progression and supporting earlier prognostic assessment[64,65].
From a data-sharing perspective, federated learning has emerged as a promising adjacent strategy for addressing the data scarcity bottleneck in HCC research. By enabling multi-institutional model training without sharing raw patient data, federated learning preserves privacy while leveraging diverse and heterogeneous datasets. Recent studies have demonstrated the feasibility of federated approaches for liver cancer CT diagnosis and outcome prediction in HCC patients undergoing radiotherapy, although federated learning should be viewed as complementary to, rather than synonymous with, MTL[32]. The development of standardized reporting frameworks, such as Transparent Reporting of a multivariable prediction model for individual prognosis or diagnosis + AI, and the integration of explainable AI (XAI) methods including SHapley Additive exPlanations, Local Interpretable Model-agnostic Explanations, and Gradient-weighted Class Activation Mapping into MTL pipelines are expected to facilitate regulatory approval and clinical adoption of these models[26]. DL models trained on whole-slide histopathology images have also demonstrated the ability to provide spatial quantification of MVI, which is essential for accurate prediction of post-surgical recurrence risk[59].
An important caveat is that not all task combinations are beneficial. When auxiliary tasks bear little biological or clinical relationship to the primary objective - for example, pairing HCC subtype classification with unrelated metabolic biomarker prediction - the shared feature space may become contaminated with irrelevant signals, resulting in negative transfer. Existing studies have demonstrated strong multitask performance in carefully curated task sets, yet performance gaps persist in specific subtasks, underscoring the need for systematic task-compatibility analysis and for benchmarking frameworks that can guide task selection[35].
From a translational perspective, the preoperative prediction of MVI through MTL has immediate clinical relevance for surgical planning. Because MVI is one of the strongest predictors of early recurrence and poor outcome after curative resection, non-invasive identification of this feature could inform decisions about resection margins, the need for anatomic vs non-anatomic hepatectomy, and candidacy for adjuvant therapy. Joint prediction of MVI and VETC within a single model further enhances risk stratification and supports personalized treatment planning[49]. Accurate preoperative assessment of both MVI status and RFS is essential for personalized HCC management. The multitask model developed by Wang et al[51], which jointly predicts MVI and RFS from preoperative MRI, demonstrated that integrating these two endpoints within a single framework improves predictive accuracy for both. Future prospective studies should determine whether such models can be incorporated meaningfully into existing staging systems and treatment algorithms.
CONCLUSION
MTL has emerged as a promising and evolving framework for HCC-oriented prediction, segmentation, prognostication, and treatment-response assessment, with a growing body of early evidence supporting its translational potential. By jointly optimizing related clinical objectives within a shared representation, MTL models may improve generalization and data efficiency compared with their single-task counterparts. The evidence reviewed here demonstrates promising findings across HCC subtype identification, preoperative MVI and VETC prediction, recurrence-risk estimation, and treatment-response assessment. At the same time, the available literature makes clear that MTL is most effective when task combinations are biologically coherent, architecturally well designed, and clinically interpretable.
Nonetheless, challenges related to task selection, loss balancing, negative transfer, and model interpretability must be addressed before widespread clinical deployment becomes feasible. Looking forward, the integration of multi-modal data, including imaging, genomics, and clinical records, advances in self-supervised and uncertainty-aware training, and the development of explainability tools tailored to multitask architectures are expected to drive further progress. Ultimately, if validated through rigorous prospective and multicenter studies, MTL may meaningfully support precision oncology in HCC by providing clinicians with integrated diagnostic and prognostic decision-support tools. Recent developments in transformer-based architectures, federated learning for privacy-preserving multi-center collaboration, and whole-slide image analysis for spatial MVI quantification further reinforce the translational potential of MTL in HCC management.
Omata M, Cheng AL, Kokudo N, Kudo M, Lee JM, Jia J, Tateishi R, Han KH, Chawla YK, Shiina S, Jafri W, Payawal DA, Ohki T, Ogasawara S, Chen PJ, Lesmana CRA, Lesmana LA, Gani RA, Obi S, Dokmeci AK, Sarin SK. Asia-Pacific clinical practice guidelines on the management of hepatocellular carcinoma: a 2017 update.Hepatol Int. 2017;11:317-370.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Cited by in Crossref: 1808][Cited by in RCA: 1787][Article Influence: 198.6][Reference Citation Analysis (16)]
Song J, Dong H, Chen Y, Zhang X, Zhan G, Jain RK, Chen YW. Early Recurrence Prediction of Hepatocellular Carcinoma Using Deep Learning Frameworks with Multi-Task Pre-Training.Information. 2024;15:493.
[PubMed] [DOI] [Full Text]
Reig M, Sanduzzi-Zamparelli M, Forner A, Rimola J, Ferrer-Fàbrega J, Burrel M, Garcia-Criado Á, Díaz A, Llarch N, Iserte G, Mollà M, Kelley RK, Galle PR, Mazzaferro V, Salem R, Sangro B, Singal AG, Vogel A, Yanagihara TK, Ayuso C, Torres F, Bruix J. BCLC strategy for prognosis prediction and treatment recommendations: The 2026 update.J Hepatol. 2026;84:631-654.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 39][Cited by in RCA: 134][Article Influence: 134.0][Reference Citation Analysis (9)]
Pfahl EL, Pracha NS, Emlemdi MH, Le PD, Makary MS. Artificial Intelligence in the Diagnosis and Prognostic Stratification of Hepatocellular Carcinoma: Current Evidence, Clinical Applications, and Future Perspectives.Biomedicines. 2026;14:505.
[RCA] [PubMed] [DOI] [Full Text] [Full Text (PDF)][Reference Citation Analysis (0)]
Xiao L, Wang J, Cui H, Zhu H, He J, Deng H, Zhang W, Dong H, Zhou Y, Jiang P, Zeng L, Peng J, Xu P, Shen R, Kurban N, Lin M, Lu S, Weng X, Hong C, Liu L. Multi-modal gradual fusion transformer-based model for predicting immunotherapy response in patients with hepatocellular carcinoma.J Adv Res. 2026;S2090-1232(26)00113.
[RCA] [PubMed] [DOI] [Full Text][Reference Citation Analysis (0)]
Nolte J, Guichelaar MMJ, Bouman DE, van den Berg SM, Haeri MA.
A CNN-Transformer for Classification of Longitudinal 3D MRI Images -- A Case Study on Hepatocellular Carcinoma Prediction. 2025 Preprint. Available from: arXiv:2501.10733v2.
[PubMed] [DOI] [Full Text]
Famularo S, Penzo C, Maino C, Milana F, Oliva R, Marescaux J, Diana M, Romano F, Giuliante F, Ardito F, Grazi GL, Donadon M, Torzilli G. Preoperative detection of hepatocellular carcinoma's microvascular invasion on CT-scan by machine learning and radiomics: A preliminary analysis.Eur J Surg Oncol. 2025;51:108274.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 12][Cited by in RCA: 11][Article Influence: 11.0][Reference Citation Analysis (0)]
Xie Y, Li S, Li X, Liu B, Xu Y, Zhou W.
Triplet-Uncertainty in Multi-Task Deep Learning for Improving Malignancy Characterization of Hepatocellular Carcinoma. 2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI); 2023 Apr 18-21; Cartagena, Colombia. Hoboken (NJ): IEEE, 2023.
[PubMed] [DOI] [Full Text]
Li Y, Zhao Y, Wang M, Li F, Chen J, Luo Y, Feng S, Lin X, Huang B.
Integrating with Segmentation by Using Multi-Task Learning Improves Classification Performance in Medical Image Analysis. 2022 IEEE 35th International Symposium on Computer-Based Medical Systems (CBMS); 2022 Jul 21-23; Shenzen, China. Hoboken (NJ): IEEE, 2022: 351-354.
[PubMed] [DOI] [Full Text]
Wang F, Zhan G, Chen QQ, Xu HY, Cao D, Zhang YY, Li YH, Zhang CJ, Jin Y, Ji WB, Ma JB, Yang YJ, Zhou W, Peng ZY, Liang X, Deng LP, Lin LF, Chen YW, Hu HJ. Multitask deep learning for prediction of microvascular invasion and recurrence-free survival in hepatocellular carcinoma based on MRI images.Liver Int. 2024;44:1351-1362.
[RCA] [PubMed] [DOI] [Full Text][Cited by in Crossref: 2][Cited by in RCA: 34][Article Influence: 17.0][Reference Citation Analysis (1)]
Wang S, Shao M, Fu Y, Zhao R, Xing Y, Zhang L, Xu Y. Deep learning models for predicting the survival of patients with hepatocellular carcinoma based on a surveillance, epidemiology, and end results (SEER) database analysis.Sci Rep. 2024;14:13232.
[RCA] [PubMed] [DOI] [Full Text][Cited by in RCA: 12][Reference Citation Analysis (0)]
Zhao Y, Huang X, Sun M, Chen J, Zhang J, Feng S, Li J, Cao K, Wang J, Huang B, Zou Y. Predicting microvascular invasion plus cytokeratin 19 expression positivity in hepatocellular carcinoma based on EOB-MRI using multitask deep learning.Hepatoma Res. 2025;11:12.
[RCA] [PubMed] [DOI] [Full Text][Cited by in RCA: 1][Reference Citation Analysis (0)]
Huang H, Liu B, Zhang L, Xu Y, Zhou W.
Transformer Based Multi-task Deep Learning with Intravoxel Incoherent Motion Model Fitting for Microvascular Invasion Prediction of Hepatocellular Carcinoma. In: Wang L, Dou Q, Fletcher PT, Speidel S, Li S, editors. Medical Image Computing and Computer Assisted Intervention – MICCAI 2022 Sep 18-22; Singapore. Cham: Springer, 2022: 266-275.
[PubMed] [DOI] [Full Text]
Wu X, Feng Y, Xu H, Lin Z, Chen T, Li S, Qiu S, Liu Q, Ma Y, Zhang S. CTransCNN: Combining transformer and CNN in multilabel medical image classification.Knowledge-Based Syst. 2023;281:111030.
[PubMed] [DOI] [Full Text]
Deshpande A, Gupta D, Bhurane A, Meshram N, Singh S, Radeva P.
Hybrid deep learning-based strategy for the hepatocellular carcinoma cancer grade classification of H&E stained liver histopathology images. 2025 Preprint. Available from: arXiv:2412.03084v2.
[PubMed] [DOI] [Full Text]
Jana A, Arunachalam R, Minacapelli CD, Catalano K, Catalano C, Rustgi V, Metaxas D.
Scale-Aware Multi-Instance Learning for Early Prognosis of Subjects at Risk of Developing Hepatocellular Carcinoma. 2023 IEEE 20th International Symposium on Biomedical Imaging (ISBI); 2023 Apr 18-21; Cartagena, Colombia. Hoboken (NJ): IEEE, 2023.
[PubMed] [DOI] [Full Text]
Footnotes
Peer review: Externally peer reviewed.
Peer-review model: Single blind
Specialty type: Oncology
Country of origin: Türkiye
Peer-review report’s classification
Scientific quality: Grade C, Grade C, Grade C
Novelty: Grade C, Grade C, Grade C
Creativity or innovation: Grade B, Grade C, Grade C
Scientific significance: Grade B, Grade C, Grade C
P-Reviewer: Yang WY, MD, China; Ye J, Academic Fellow, China S-Editor: Zuo Q L-Editor: A P-Editor: Zhao S