Abstract
Quantifying pulmonary vascular remodeling from histology still relies heavily on manual vessel delineation, which is labor-intensive and difficult to scale. Although modern segmentation backbones produce strong vessel masks, their errors are often complementary. We therefore formulate the task as posterior fusion and introduce ReliFuse, which combines cached expert probability maps without accessing the RGB image at the fusion stage. ReliFuse constructs ensemble-state features, estimates validation-anchored local reliability, pools evidence in logit space, and applies bounded residual corrections only in ambiguous regions. The fusion head is trained with overlap, boundary, consensus-preservation, sparse-correction, and calibration objectives to refine uncertain areas while preserving confident agreement. On a public pulmonary histology dataset, ReliFuse achieves the highest primary overlap in the matched posterior-stack comparison. Although its Dice and IoU gains over the strongest learned fusion heads are small, improvements are clearest in disagreement-heavy and vessel-rich cases. These findings support reliability-calibrated posterior fusion as a practical approach when multiple frozen experts are available.
Data Availability
The dataset used in this study is publicly available. The microscopy images and corresponding expert-annotated segmentation masks can be accessed from the dataset published by Sinitca et al. Sinitca, A.M., Lyanova, A.I., Kaplun, D.I., Hassan, H., Krasichkov, A.S., Sanarova, K.E., Shilenko, L.A., Sidorova, E.E., Akhmetova, A.A., Vaulina, D.D., et al.: Microscopy image dataset for deep learning-based quantitative assessment of pulmonary vascular changes. Scientific Data 11(1), 635 (2024).
Code Availability
The source code for the ReliFuse framework is available at https://github.com/letruongzzio/ReliFuse
References
Ba, X., Zhang, X., Li, S., Yuan, J., & Hu, J. (2025). Multimodal semantic communication system based on graph neural networks. Intelligence & Robotics,5(3), 805–826. https://doi.org/10.20517/ir.2025.41
Chen, L.-C., Zhu, Y., Papandreou, G., Schroff, F., & Adam, H. (2018). Encoder-decoder with atrous separable convolution for semantic image segmentation. Proceedings of the European Conference on Computer Vision (ECCV) (pp. 801–818)
Dietterich, T. G. (2000). Ensemble methods in machine learning. International Workshop on Multiple Classifier Systems (pp. 1–15). Springer.
Frangi, A.F., Niessen, W.J., Vincken, K.L., & Viergever, M.A. (1998). Multiscale vessel enhancement filtering. In: Medical Image Computing and Computer-Assisted Intervention—MICCAI’98, pp. 130–137. Springer
Guo, C., Pleiss, G., Sun, Y., & Weinberger, K.Q. (2017). On calibration of modern neural networks. In: Precup, D., Teh, Y.W. (eds.) Proceedings of the 34th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 70, pp. 1321–1330. PMLR, Sydney, Australia. https://proceedings.mlr.press/v70/guo17a.html
Huang, H., Lin, L., Tong, R., Hu, H., Zhang, Q., Iwamoto, Y., Han, X., Chen, Y.-W., & Wu, J. (2020). Unet 3+: A full-scale connected unet for medical image segmentation. In: ICASSP 2020 – IEEE International Conference on Acoustics, Speech and Signal Processing, pp. 1055–1059. IEEE
Lakshminarayanan, B., Pritzel, A., & Blundell, C. (2017). Simple and scalable predictive uncertainty estimation using deep ensembles. Proceedings of the 31st International Conference on Neural Information Processing Systems. NIPS’17 (pp. 6405–6416). Red Hook, NY, USA: Curran Associates Inc.
Lin, T.-Y., Dollár, P., Girshick, R., He, K., Hariharan, B., & Belongie, S. (2017). Feature pyramid networks for object detection. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 2117–2125)
Litjens, G., Kooi, T., Bejnordi, B. E., Setio, A. A. A., Ciompi, F., Ghafoorian, M., Laak, J. A., Ginneken, B., & Sánchez, C. I. (2017). A survey on deep learning in medical image analysis. Medical Image Analysis,42, 60–88.
Liu, Z., Mao, H., Wu, C.-Y., Feichtenhofer, C., Darrell, T., & Xie, S. (2022). A convnet for the 2020s. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 2968–2979)
Liu, D., Zhang, Y., Ren, Y., Pan, D., He, X., Cong, M., & Yu, G. (2025). Multi-directional attention: A lightweight attention module for slender structures. Intelligence & Robotics 5(4), 827–843. 10.20517/ir.2025.42
Long, J., Shelhamer, E., & Darrell, T. (2015). Fully convolutional networks for semantic segmentation. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3431–3440. https://doi.org/10.1109/CVPR.2015.7298965
Ma, J., Li, F., & Wang, B. (2024). U-Mamba: Enhancing long-range dependency for biomedical image segmentation. arXiv preprint arXiv:2401.04722. https://doi.org/10.48550/arXiv.2401.04722
Mehta, S., & Rastegari, M. (2022). MobileViT: Light-weight, general-purpose, and mobile-friendly vision transformer. International Conference on Learning Representations
Mirikharaji, Z., Abhishek, K., Izadi, S., Hamarneh, G.: D-LEMA: Deep learning ensembles from multiple annotations – application to skin lesion segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 1837–1846 (2021). 10.1109/CVPRW53098.2021.00203 .
Moccia, S., De Momi, E., El Hadji, S., & Mattos, L. S. (2018). Blood vessel segmentation algorithms—review of methods, datasets and evaluation metrics. Computer Methods and Programs in Biomedicine,158, 71–91.
Ni, J., Shi, J., Zhan, Q., Zhang, Z., & Gu, Y. (2025). An improved u-net model with multiscale fusion for retinal vessel segmentation. Intelligence & Robotics,5(3), 679–694. https://doi.org/10.20517/ir.2025.35
Oktay, O., Schlemper, J., Folgoc, L.L., Lee, M.C.H., Heinrich, M.P., Misawa, K., Mori, K., McDonagh, S.G., Hammerla, N.Y., Kainz, B., Glocker, B., & Rueckert, D. (2018). Attention u-net: Learning where to look for the pancreas. In: Medical Imaging with Deep Learning. MIDL 2018 Oral. https://openreview.net/forum?id=Skft7cijM
Ronneberger, O., Fischer, P., & Brox, T. (2015). U-net: Convolutional networks for biomedical image segmentation. In N. Navab, J. Hornegger, W. M. Wells, & A. F. Frangi (Eds.), Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015 (pp. 234–241). Cham: Springer.
Sinitca, A. M., Lyanova, A. I., Kaplun, D. I., Hassan, H., Krasichkov, A. S., Sanarova, K. E., Shilenko, L. A., Sidorova, E. E., Akhmetova, A. A., Vaulina, D. D., et al. (2024). Microscopy image dataset for deep learning-based quantitative assessment of pulmonary vascular changes. Scientific Data,11(1), 635.
Tan, M., Le, Q.V.: EfficientNetV2: Smaller models and faster training. In: Proceedings of the 38th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 139, pp. 10096–10106. PMLR, Virtual Event (2021). https://proceedings.mlr.press/v139/tan21a.html
Tetteh, G., Efremov, V., Forkert, N. D., Schneider, M., Kirschke, J., Weber, B., Zimmer, C., Piraud, M., & Menze, B. H. (2020). Deepvesselnet: Vessel segmentation, centerline prediction, and bifurcation detection in 3-d angiographic volumes. Frontiers in Neuroscience,14, Article 592352.
Tu, Z., Talebi, H., Zhang, H., Yang, F., Milanfar, P., Bovik, A., & Li, Y. (2022). MaxViT: Multi-axis vision transformer. European Conference on Computer Vision (pp. 459–479). Springer.
Wang, J., Jin, Y., Wang, L. (2022). Personalizing federated medical image segmentation via local calibration. arXiv preprint arXiv:2207.04655. 10.48550/arXiv.2207.04655
Wang, Y., Luo, L., Wu, M., Wang, Q., & Chen, H. (2025). Learning robust medical image segmentation from multi-source annotations. Medical Image Analysis 101, 103489. 10.1016/j.media.2025.103489
Wolpert, D. H. (1992). Stacked generalization. Neural Networks,5(2), 241–259. https://doi.org/10.1016/S0893-6080(05)80023-1
Xiao, T., Liu, Y., Zhou, B., Jiang, Y., & Sun, J. (2018). Unified perceptual parsing for scene understanding. Proceedings of the European Conference on Computer Vision (ECCV) (pp. 418–434)
Xie, E., Wang, W., Yu, Z., Anandkumar, A., Alvarez, J. M., & Luo, P. (2021). Segformer: Simple and efficient design for semantic segmentation with transformers. Advances in Neural Information Processing Systems (Vol. 34, pp. 12077–12090)
Yu, W., Si, C., Zhou, P., Luo, M., Zhou, Y., Feng, J., Yan, S., & Wang, X. (2024). Metaformer baselines for vision. IEEE Transactions on Pattern Analysis and Machine Intelligence,46(2), 896–912. https://doi.org/10.1109/TPAMI.2023.3329173
Zhou, Z., Rahman Siddiquee, M. M., Tajbakhsh, N., & Liang, J. (2018). Unet++: A nested u-net architecture for medical image segmentation. In D. Stoyanov, Z. Taylor, G. Carneiro, T. Syeda-Mahmood, A. Martel, L. Maier-Hein, J. M. R. S. Tavares, A. Bradley, J. P. Papa, V. Belagiannis, J. C. Nascimento, Z. Lu, S. Conjeti, M. Moradi, H. Greenspan, & A. Madabhushi (Eds.), Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support (pp. 3–11). Cham: Springer.
Acknowledgements
The authors wish to extend their sincere thanks to Aleksandr Sinitca and colleagues for creating and publicly sharing the Microscopy Image Dataset for Pulmonary Vascular Changes, which made this research possible.
Funding
Son T. Huynh was funded by the PhD Scholarship Programme of Vingroup Innovation Foundation (VINIF), code VINIF.2025.TS.02.
Author information
Authors and Affiliations
Contributions
- Truong P. Le: contributed to conceptualization, methodology, software, and writing of the original draft. - Truong N. Nguyen: contributed to data analysis, data preprocessing, methodology, software, visualization, and writing of the original draft. - Vinh Le H. Tran: contributed to software, validation, and writing-review and editing - Tram T. Doan & Binh T. Nguyen: contributed to validation, review, and supervision. - Son T. Huynh: contributed to conceptualization, supervision, project administration, and writing-review and editing.
Corresponding author
Ethics declarations
Conflict of interest
The authors declare that they have no competing financial interests or personal relationships that could have influenced the work reported in this paper.
Ethical Approval and Consent to Participate
This study is based entirely on an existing, publicly available dataset of microscopy images from rat models of pulmonary hypertension, originally published by Sinitca et al. (Sinitca et al., 2024). This article does not contain any studies with human participants or animals performed by any of the authors. Therefore, no additional ethics approval or informed consent was required.
Consent for Publication
Not applicable as the manuscript does not contain any individual person’s data.
Additional information
Editors: Gianvito Pio, Jurica Levatić, Nikola Simidjievski.
Publisher's Note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Supplementary Information
Below is the link to the electronic supplementary material.
Appendices
Appendix A: Supporting Expert-Bank Construction
For reproducibility, we give the full procedure used to construct the experimental expert bank. Algorithm 1 implements the admissibility screen associated with Eq. (5), and Algorithm 2 gives the greedy approximation to Eq. (4). We then examine how this bank choice affects several fusion rules in Table 5. These details support the ReliFuse experiments; they are not presented as a separate model-selection contribution.
Admissible Candidate Pool. The complete Stage-A outcome appears in Table 4. Its purpose is to show that the seven-expert stack was drawn from stable segmenters with broadly comparable validation Dice but different precision–recall profiles and capacities. We retain the audit here because the primary claim concerns posterior fusion, not image-to-mask model selection.
Supporting Expert-Bank Sensitivity. We next ask whether the fusion result is sensitive to this supporting choice. The Top-7 bank contains the seven experts with the highest validation Dice, whereas the diversity-aware bank follows Sect. 3.2. The former has a higher mean base Dice (0.9030 versus 0.8958), but substantially less pairwise disagreement (0.0427 versus 0.0757). Table 5 compares the two banks within this dataset; it does not imply that either construction rule is universally preferable.
In this dataset, the diversity-aware bank improves every matched rule, including averaging, voting, and ReliFuse. For ReliFuse specifically, Dice rises by 0.0043, IoU by 0.0069, and recall from \(0.9219\,{\scriptstyle \pm \,0.0317}\) to \(0.9297\,{\scriptstyle \pm \,0.0188}\). The comparison confirms that bank composition can influence posterior fusion and should therefore be documented. It does not alter the methodological scope: ReliFuse is the proposed method, while the bank construction remains supporting experimental setup.
Appendix B: ReliFuse Ablation Studies
The primary results use the lightweight boundary-aware ReliFuse configuration. To explain that choice, we first compare internal variants that share the posterior input in Sect. 3.3.1 and the loss family in Eq. (32), but vary the ambiguity gate, residual strength, local calibration, and boundary emphasis. Table 6 is therefore a design screen, not an all-method benchmark.
The boundary-aware configuration was selected using the fusion-validation criterion and its conservative correction profile before final held-out reporting. Table 6 reports held-out behavior post hoc; it was not used to choose the variant. In this held-out screen, the high-sensitivity variant has the highest mean recall, but its precision falls from \(0.9194\,{\scriptstyle \pm \,0.0299}\) to \(0.9067\,{\scriptstyle \pm \,0.0312}\) and its correction magnitude increases to \(|\Delta |=2.805\pm 0.035\). Strong calibration moves in the opposite direction, reducing ambiguity and correction scale to \((U=0.082\pm 0.005,\ |\Delta |=0.854\pm 0.006)\), but also lowering Dice and IoU. The selected boundary-aware configuration gives the best mean Dice and IoU in the post hoc screen, ties for the best Hard Dice, and retains a moderate correction profile \((U=0.151\pm 0.004,\ |\Delta |=1.339\pm 0.015)\).
Having selected the configuration, we separately remove or simplify individual non-image components in Table 7. This second analysis is intended to reveal mechanism sensitivity rather than to repeat variant selection. We include only changes that remain meaningful under the current method definition, focusing on whether calibrated evidence, boundary-aware ambiguity, and constrained correction contribute beyond generic stacker capacity.
The component screen complements rather than replaces the selected result in Table 3. Without learned ambiguity refinement, Dice is lowest and the source diagnostics show very small ambiguity and residual magnitudes (\(U=0.013\), \(|\Delta |=0.293\)), suggesting that the fixed scaffold alone is too restrictive. Removing the gate has the opposite effect: U becomes one everywhere, permitting dense correction and weakening the consensus-preservation principle in Eq. (26). The loss ablations remain close in overlap, yet the calibration, boundary, and consensus terms keep the objective aligned with the failure modes the method is designed to address.
Appendix C: Extended Comparisons and Stress Tests
The main table emphasizes overlap and foreground classification. Here we ask two narrower questions: whether those gains are compatible with reasonable boundary geometry, and whether they persist under difficult subsets. Table 8 therefore adds boundary and distance diagnostics for the same method families, as supporting evidence rather than new primary endpoints.
ReliFuse remains close to the strongest boundary scores, so its Dice and IoU result is not accompanied by a collapse in contour quality. It is not best on every measure: P-MoLE has the highest boundary F1, precision, and recall, while D-LEMA is strongest on HD95 and ASSD. These metrics thus complement the primary endpoint and help delimit the claim rather than simply reproducing the same ranking.
We then isolate stress groups that more directly probe the proposed mechanism. Each column in Table 9 contains the eight held-out batches at the relevant extreme of disagreement, vessel area, weighted Dice, or weighted recall. The latter two groups are especially useful because they test whether a learned fusion method repairs cases in which conventional weighted averaging is already weak.
The strongest ReliFuse results occur in the high-disagreement and high-vessel-area groups, which are also the conditions most closely aligned with its ambiguity and minority-evidence mechanisms. On the low weighted-Dice and low weighted-recall subsets, ReliFuse remains competitive but is not always best; P-MoLE, D-LEMA, and the shallow stacker also handle parts of these broader failure regimes. Reporting the complete matrix makes this boundary explicit and avoids treating the two favorable subsets as the whole stress-test result.
The paired analysis reinforces a cautious reading. Dice and IoU deltas are positive against every listed reference, with clearer separation from fixed rules, the shallow stacker, and UMA-Net. Against D-LEMA, LC-Fed, and P-MoLE, however, the effects are small and not significant for Dice or IoU. We therefore treat those rows as evidence about effect direction and magnitude, not as proof of broad superiority. The recall deltas likewise expose the expected trade-off with recall-heavy stackers.
Appendix D: Advanced Backbone and Source-Bank References
The primary benchmark fixes one seven-expert bank, so we add a smaller source-bank check with stronger contemporary sources. In Table 11, each direct reference consumes an RGB image, whereas the corresponding “+ReliFuse” setting combines cached maps from a completed external posterior-bank run. We keep this analysis outside the matched benchmark because direct single-backbone inference and source-bank fusion are not the same computational object, and because the comparison does not isolate the contribution of ReliFuse from the contribution of multi-source aggregation.
The table is consequently organized as a direct-versus-fused check rather than as a single ranking. Direct rows reproduce the completed raw-image results, fused rows come from the corresponding external source-bank evaluation, and all rows use the same held-out protocol with 23 batch-level units. The source-bank references were chosen to cover complementary modern design choices without turning the paragraph into a single citation list: MaxViT, EfficientNetV2, and ConvNeXt represent stronger backbone families (Tu et al., 2022; Liu et al., 2022; Tan and Le, 2021); UPerNet, CAFormer/MetaFormer, and MobileViT provide contextual segmentation or lightweight hybrid sources (Xiao et al., 2018; Yu et al., 2024; Mehta and Rastegari, 2022); and FPN and U-Net++ supply the remaining pyramid and nested-decoder components (Zhou et al., 2018; Lin et al., 2017). We omit parameter counts because they would compare a single segmentation network with an entire source bank plus fusion head and would therefore be misleading.
The external source-bank pipelines that incorporate ReliFuse have higher Dice and IoU than their listed direct raw-image references, with the largest change for ConvNeXt-Tiny + UPerNet and only a small change for the already strong MaxViT-Tiny + U-Net++ reference. This should be read as contextual evidence, not as a causal with-versus-without ReliFuse test on a matched posterior bank. It does not expand the scope of the controlled benchmark in Table 3, and the evidence remains preliminary: a stronger generality claim would require several independently constructed banks and matched comparisons among non-ReliFuse fusion rules and ReliFuse on each bank.
Appendix E: Cached-Fusion Efficiency
Because the central experiments concern segmentation quality, we report computational cost separately. Every row in Table 12 starts from cached predictions and therefore measures only the fusion stage. The compared heads include fixed aggregation, stackers based on stacked generalization (Wolpert, 1992), and learned posterior-fusion references including D-LEMA, LC-Fed, and UMA-Net (Mirikharaji et al., 2021; Wang et al., 2022, 2025); P-MoLE is included as the implemented posterior mixture-of-local-experts comparator described in the experimental protocol. For context, recomputing the seven raw-image experts in the same profiling setup requires approximately 9356.61 ms per batch of four images and 13404.11 MB of peak allocated memory, far more than any cached-fusion pass.
ReliFuse is slower than fixed cached rules because it must construct the diagnostic state, estimate calibrated opinions, and apply ambiguity-gated correction. Its parameter count nevertheless remains modest relative to the base experts. FLOPs and MACs are omitted because the profiler did not support all custom operators; these quantities should be regenerated when a compatible profiling environment is available.
Appendix F: Calibration and Morphology Diagnostics
Since ReliFuse explicitly models reliability and includes the calibration term in Eq. (31), overlap alone is not sufficient to characterize its output. Table 13 pairs global calibration measures with vessel-oriented morphology diagnostics for the strongest reference families. Boundary and distance measures already appear in Table 8 and are not repeated.
No method dominates this diagnostic view. LC-Fed has the lowest calibration errors, D-LEMA the highest clDice, and P-MoLE the lowest false-connection proxy. ReliFuse remains close on the morphology measures while producing the strongest primary Dice and IoU in Table 3. Its global calibration scores should also be read alongside the diagnostic maps, because the method is designed to calibrate expert evidence locally rather than to optimize only a whole-test-set calibration statistic. Exact definitions for these auxiliary metrics are given in Appendix H.3.
Appendix G: Additional Qualitative Diagnostics
To connect the equations with observable behavior, Fig. 5 follows the calibrated prior p, ambiguity field U, gated residual, and final output through several held-out examples. These cases extend the compact comparison in Fig. 3 and include both successful corrections and a failure.
In the limitation case, ReliFuse recovers additional foreground but loses enough precision to reduce Dice relative to validation-weighted averaging. The same gate that can recover a faint or disputed vessel can therefore enlarge a false-positive region when the posterior evidence is misleading. By contrast, low-U regions remain close to the prior, illustrating the consensus-preservation behavior encouraged by Eq. (29).
Appendix H: Implementation and Reproducibility Details
We close the Appendix by recording how the reported results were produced. The lightweight boundary-aware ReliFuse configuration was selected by the fusion-validation criterion in the configuration-screening study. The ReliFuse row in Table 3, the paired statistics, and the all-method boundary, hard-subset, efficiency, calibration, morphology, and distance results come from the common-protocol all-method comparison using that selected configuration. Table 15 makes these sources explicit.
1.1 Fusion-Head Architecture
The selected ReliFuse configuration uses three fully convolutional branches and preserves the spatial resolution of every posterior map. For compactness, let \(\mathcal {C}_{d}^{a\rightarrow h}\) denote a bias-free \(3\times 3\) convolution from a to h channels with dilation d and same padding, followed by GroupNorm and SiLU. The number of groups is the largest divisor of h in \(\{8,4,2,1\}\). We also use a residual block \(\mathcal {B}_{d}^{h}\) consisting of \(\mathcal {C}_{d}^{h\rightarrow h}\), a second bias-free \(3\times 3\) dilated convolution with GroupNorm, residual addition, and SiLU activation. No pooling, upsampling, dropout, or RGB-image pathway is used in the fusion head.
Table 14 gives the exact branch specification for the reported \(K=7\) configuration. Its diagnostic state has ten channels, including the support-gap proxy G defined above, and the common hidden width is \(h=24\). The ambiguity branch uses \(h_A=12\). The reliability branch \(H_{\theta }\) returns \(2K=14\) maps, which are split channel-wise into the raw reliability logits \(a_i\) and calibration-bias logits \(c_i\). By contrast, \(A_{\theta }\) and \(R_{\theta }\) each return one map for the ambiguity refinement and residual correction, respectively.
The selected boundary-aware variant sets \(s_b=1.25\), \(\lambda _U=0.20\), and \(\Delta _{\max }=2.0\) and does not include an additional topology-repair branch. We initialize every terminal \(1\times 1\) projection to zero. Consequently, the initial bias terms vanish, the ambiguity field begins at its deterministic scaffold \(\bar{U}\), and the residual correction begins at zero. This initialization makes the initial output follow the reliability-anchored log-opinion prior rather than an arbitrary learned correction. The validation priors \(\pi _i\) are registered as fixed buffers and are therefore excluded from the trainable parameter count.
1.2 Training Objective Details
This subsection makes explicit how Eq. (32) trains the fusion head. The segmentation loss uses binary cross-entropy and soft Dice on \(p_{\textrm{out}}\):
BCE supplies dense probability supervision, while Dice counteracts foreground imbalance by optimizing overlap. Both terms update the reliability, ambiguity, and residual branches through the final output.
The auxiliary terms determine where those updates are allowed to concentrate. The boundary loss uses the same Sobel-gradient operator as the diagnostic boundary cue and compares gradient magnitudes between prediction and mask, which makes contour errors visible even when the foreground area is small. The consensus loss is intentionally asymmetric: \(\operatorname {sg}(p)\) prevents the prior from being pulled toward the final output, and \(\operatorname {sg}(U)\) prevents the loss from lowering the gate simply to avoid a penalty. As a result, low-ambiguity pixels train the residual path to leave the prior unchanged, whereas high-ambiguity pixels remain available for correction.
The sparse loss is applied to \(U\Delta \), the correction actually added to the logit. This is important because a large raw residual in a low-ambiguity region has little effect, while a sustained gated correction is penalized. The calibration loss supervises both the prior and the final posterior with a Brier-style mean-squared error. In practice, this keeps the log-opinion pool useful before residual correction and reduces the incentive for the residual branch to compensate for a poorly calibrated prior. The coefficients in Eq. (32) were treated as part of the validation-screened ReliFuse configuration and were fixed before the common-protocol held-out comparison.
1.3 Metric definitions
All hard-mask metrics use a fixed threshold of 0.5 for every method. For a predicted binary mask P and ground truth G, the batch-level overlap metrics accumulate true positives, false positives, false negatives, and true negatives over all images in the evaluation batch before applying
with \(\epsilon =10^{-6}\). This is the protocol used for Table 3, the hard-subset matrices, and the paired batch-level statistics. The hard subsets in Table 9 are formed from the same 23 evaluation batches: high-disagreement and high-vessel-area select the eight largest batches by the corresponding stress variable, while low weighted-Dice and low weighted-recall select the eight weakest batches under the validation-weighted ensemble.
Boundary and distance metrics are computed per image and then averaged over the held-out batches. The object boundary is the one-pixel mask difference between a binary mask and its one-step binary erosion. Boundary precision and boundary recall use a tolerance radius of two pixels:
HD95 is the 95th percentile of the symmetric boundary-to-boundary Euclidean distance set, and ASSD is the mean of the same symmetric distance set. If either boundary is empty, HD95 and ASSD are recorded as missing for that image and skipped in the batch mean.
For topology-oriented morphology, skeletons are computed with two-dimensional morphological skeletonization. If that implementation is unavailable, the fallback is the one-pixel erosion boundary approximation used by the evaluation script. Let \(\operatorname {Skel}(\cdot )\) denote this skeleton operator. We report
The reported thin-vessel recall is this skeleton-recall proxy \(T_{\textrm{rec}}\); no local-radius threshold is applied. Thus it measures recovery of centerline-like ground-truth vessel structure, including thin branches, rather than a separately radius-stratified subset. If the ground-truth skeleton is empty, the \(\epsilon \) convention assigns recall one when no centerline target exists.
The false-connection proxy is component based. Connected components are counted with two-dimensional four-connectivity, yielding \(C_P\) predicted components and \(C_G\) ground-truth components. We use
This quantity penalizes predicted masks that merge multiple ground-truth components into fewer connected structures. It is not a complete false-positive measure; isolated spurious components are instead reflected by precision, Dice, IoU, boundary metrics, and component-count error. When \(C_G=0\), the denominator is one and the proxy is zero by definition.
Calibration is estimated on up to \(5\times 10^5\) sampled held-out pixels using ten equal-width probability bins. Probabilities are clipped to \([10^{-6},1-10^{-6}]\). For bin \(B_m\), let \(\bar{p}_m\) be the mean predicted probability, \(\bar{y}_m\) the observed foreground frequency, and \(n_m\) the number of sampled pixels in that bin. The expected calibration error is
Brier score is \(N^{-1}\sum _j(p_j-y_j)^2\), and NLL is
Efficiency metrics in Table 12 are measured on cached prediction stacks with batch size four. Latency is wall-clock fusion time per batch, throughput is images per second derived from that latency, peak memory is peak allocated accelerator memory during the fusion pass, and the parameter count includes only the posterior-fusion module.
1.4 Dataset and Training Protocol
The 609 paired RGB microphotographs and binary masks are split into 517 development images and 92 held-out test images. For independent learned-fusion runs, the development partition is further split into 413 fusion-training images and 104 validation images for early stopping and reliability-prior estimation. Images are resized to \(512\times 512\) and normalized with ImageNet statistics. Training augmentation includes rotations up to \(\pm 20^\circ \), horizontal and vertical flips, affine scaling from 0.8 to 1.2, translation up to 10%, shear up to \(\pm 15^\circ \), color perturbation, contrast/brightness perturbation, and Gaussian blur. Base experts are trained with AdamW, learning rate \(10^{-4}\), weight decay 0.01, automatic mixed precision, batch size 8, cosine annealing, and early stopping after five epochs without validation-loss improvement.
1.5 Training Curves and Checkpoint Selection
Because ReliFuse is a learned fusion head trained on cached posterior stacks, checkpoint selection must be tied to the validation split rather than to the held-out test set. Figure 6 reports the mean training objective, validation objective, and validation Dice over three independent learned-fusion runs. For each run, the checkpoint evaluated on the held-out test set is selected by the lowest validation loss. Validation Dice is shown only as a convergence diagnostic and is not used to select the final test checkpoint.
All learned fusion heads in the matched benchmark receive the same fixed seven-expert stack. Where repeated runs are available, we report their mean and standard deviation over three independent trainings; deterministic rules are evaluated once on the same cache. Validation loss selects each learned checkpoint before held-out testing, as illustrated for ReliFuse in Fig. 6. During evaluation, we accumulate true positives, false positives, and false negatives within each batch before computing Dice, IoU, precision, and recall. This batch-level protocol is shared by Table 3 and the corresponding Appendix tables.
Rights and permissions
Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.
About this article
Cite this article
Le, T.P., Nguyen, T.N., Tran, V.L.H. et al. ReliFuse: Reliability-Calibrated Posterior Fusion for Histological Vessel Segmentation. Mach Learn 115, 218 (2026). https://doi.org/10.1007/s10994-026-07154-3
Received:
Revised:
Accepted:
Published:
Version of record:
DOI: https://doi.org/10.1007/s10994-026-07154-3
Facts Only
* Truong P. Le, Truong N. Nguyen, and colleagues developed ReliFuse.
* ReliFuse is a posterior fusion framework for quantifying pulmonary vascular remodeling from histology.
* The framework combines cached expert probability maps without accessing raw RGB images during fusion.
* The system utilizes ensemble-state features, local reliability estimation, and bounded residual corrections.
* The model was tested on a public pulmonary histology dataset originally published by Sinitca et al. (2024).
* The dataset contains 609 paired RGB microphotographs and binary masks.
* Evaluation used 92 held-out test images.
* The primary performance metric is the overlap of vessel masks.
* ReliFuse was compared against fixed rules and learned fusion heads including D-LEMA, LC-Fed, P-MoLE, and UMA-Net.
* The source code is hosted on GitHub.
* Funding was provided by the Vingroup Innovation Foundation (VINIF).
Executive Summary
ReliFuse addresses the labor-intensive nature of manual vessel delineation in pulmonary histology by fusing the outputs of multiple frozen segmentation "experts." Rather than processing raw images at the fusion stage, it operates on cached probability maps, utilizing a reliability-calibrated approach to resolve disagreements between different segmentation models. The framework identifies ambiguous regions and applies residual corrections to refine results while preserving areas of confident agreement.
Performance results indicate that ReliFuse achieves the highest primary overlap in matched posterior-stack comparisons, with the most significant improvements appearing in vessel-rich and high-disagreement cases. However, the gains in Dice and IoU metrics over the strongest existing learned fusion heads are characterized as small. While the method shows competitiveness in boundary geometry and morphology, it does not dominate all diagnostic measures, as other models like P-MoLE and D-LEMA outperform it in specific boundary and distance metrics. The efficiency of the cached-fusion approach significantly reduces computational overhead compared to recomputing raw-image experts.
Full Take
This research follows a rigorous academic design, employing a held-out test set and a systematic ablation study to justify its architectural choices. The methodology is sound, particularly the decision to isolate the fusion head's performance from the base experts' capacities. A peer reviewer would likely note the relatively small margin of improvement over top-tier competitors (P-MoLE, D-LEMA), which suggests the field may be approaching a ceiling for posterior-only fusion without reintegrating raw image context.
The authors are intellectually honest regarding their claims; they explicitly state that Dice and IoU gains are small and that the method is not the best across every morphology metric. This avoids the common academic pitfall of overclaiming "superiority" when the data only supports "competitiveness." The novelty lies not in a massive leap in accuracy, but in the reliability-calibrated mechanism and the computational efficiency of the cached-fusion pipeline.
If these findings hold, the real-world implication is a scalable workflow where practitioners can deploy a "bank" of diverse, frozen models and use a lightweight head to synthesize the best result, reducing the need for expensive re-inference. To further validate these claims, a study using diverse datasets beyond the Sinitca collection is necessary to ensure the "diversity-aware bank" construction generalizes across different histological pathologies.
Bridge Questions:
1. At what point does the loss of raw RGB data at the fusion stage become a bottleneck that no amount of posterior calibration can overcome?
2. How does the performance of a diversity-aware bank compare to a bank of experts trained on fundamentally different architectures (e.g., Transformers vs. CNNs) rather than just different training profiles?
3. Could this reliability-calibrated fusion be applied to multi-modal medical data where "experts" are different imaging modalities?
Counterstrike Scan: The content is a standard academic contribution with transparent methodology and modest claims; it does not align with any known influence campaign patterns.
