PO.BCS01.07 · 生物信息与计算

用于数字化和标准化存档泌尿生殖系统癌症切片的高通量切片扫描流程

A high-throughput slide scanning pipeline for digitizing and standardizing legacy genitourinary cancer slides

海报缩略图:用于数字化和标准化存档泌尿生殖系统癌症切片的高通量切片扫描流程
编号 1452 展板 15 时间 4/20 09:00–12:00 区域 Section 4 主讲 Faria Kabir, Unknown
分会场 Digital Pathology 2
查看 PDF 下载 PDF 🔒 查看 / 下载完整 PDF 需登录并开通下载套餐 · 查看套餐 / 开通 AACR 官方页面

作者与单位 Authors & Affiliations

Faria Kabir1, Mitra Shavakhi2, Akash Parvatikar1, Adrien Cesaire1, Egypt Phillips1, Rachel Trowbridge2, Liliana Ascione2, Pablo Barrios2, Marc Eid2, Jasmine Lee2, Sabina Signoretti3, Eliezer Van Allen2, Atish Choudhury2, Linh Hoang1, Toni K. Choueiri2, Jeremiah Wala2

1HistoWiz, Long Island City, NY,2Dana-Farber Cancer Institute, Boston, MA,3Brigham and Women’s Hospital, Boston, MA

摘要 Abstract

中文摘要
目的: 利用一家大型癌症研究机构的存档切片,为泌尿生殖系统癌症构建一个可重现的数字化切片存档流程和组织图像库,用于数字病理学应用和生物标志物发现。 背景: 数字病理学将计算机视觉应用于数字化的H&E切片,以量化肿瘤微环境和结构特征。具有广泛临床随访的存档组织对于这一目标至关重要,但将玻璃切片转换为高质量数字图像是一项重大挑战。我们描述了在HistoWiz开发的高通量全切片成像(WSI)流程及其在来自Dana-Farber癌症研究所(DFCI)Gelb转化研究中心的20,000余张存档癌症切片上的应用。 方法: HistoWiz设计了一套工作流程,包括:(i) 对染色质量、标记格式和盖玻片伪影各异的存档临床切片进行标准化接收;(ii) 专有的批量物流解决方案,使每批次可运输多达9,600张切片,损坏率<0.001%;(iii) 使用能够穿透老化切片上常见的污垢/薄膜层的硬件进行高保真扫描;(iv) 通过HistoWiz PathologyMap平台进行实时、可扩展的审查流程;以及 (v) 对手写切片标签进行光学字符识别(OCR)。 结果: 跨越20多年(2001年至2025年)收集的存档切片,使用两个高通量扫描集群以每天900张的速度进行扫描,初始QC通过率>90%。一名病理学家对图像的代表性子集进行了独立评估,评估的问题包括组织折叠、开裂、气泡、疱状物、墨水标注和失焦区域。我们发现,大多数QC失败由四类反复出现的问题引起:切片机操作导致的组织折叠(30%)、盖玻片问题(10%)、封片介质残留(2%)以及手写标注或处理伪影(1.5-12.5%)。纠正措施包括重新扫描、二甲苯或酒精基盖玻片清洁以及重新盖片,解决了>95%的盖玻片或封片介质问题病例,且组织损失有限。最终切片以金字塔化OME-TIFF文件格式存储,分辨率为0.248 µm/像素,平均每张切片2 GB。在一个离线Python流程中,一个源自Google Vision的OCR模型在组织学切片上的条形码读取率和准确性方面表现优异,优于Microsoft的OCR、Tesseract(v4+)和一个基于Keras的CRNN基线模型。 结论: 我们的大批量、高通量WSI存档工作流程提供了一个可扩展的成像流程,实现了一个20年学术组织生物库的数字化,并提供对AI就绪、高质量全分辨率切片图像的按需访问。这一数字图像库将支持大规模数字病理学研究,以发现泌尿生殖系统癌症的生物标志物。
查看英文原文 English abstract
Objective: To build a reproducible digital slide archival pipeline and tissue image repository for genitourinary cancers using archival slides from a large cancer research organization for digital pathology applications and biomarker discovery. Background: Digital pathology applies computer vision to digitized H&E to quantify tumor microenvironment and architectural features. Archival tissue with extensive clinical follow-up is essential to this aim, but converting glass slides to high-quality digital images is a significant challenge. We describe a high-throughput, whole slide imaging (WSI) pipeline developed at HistoWiz and its application to over 20,000 legacy cancer slides from the Gelb Center for Translational Research at the Dana-Farber Cancer Institute (DFCI). Methods: HistoWiz designed a workflow for (i) standardized intake of archival clinical slides with heterogeneous stain quality, labeling formats, and coverslipping artifacts, (ii) proprietary bulk logistics solutions enabling transport of up to 9,600 slides per shipment with <0.001% damage rate, (iii) high-fidelity scanning using hardware capable of penetrating dirt/film layers commonly found on aging slides, (iv) a real-time, scalable review process via the HistoWiz PathologyMap platform and (v) optical character recognition (OCR) for handwritten slide labels. Results: Archived slides from over 20 years (2001 to 2025) of slide collection were scanned at 900 slides per day, using two high-throughput scanning clusters with a >90% initial QC pass rate. A representative subset of images was independently evaluated by a pathologist for issues including tissue folds, cracking, air bubbles, blebs, ink annotations, and out-of-focus regions. We found that most QC failures were caused by four recurring issues: tissue folding from microtomy (30%), cover-slip issues (10%), mounting media residue (2%), and handwritten annotations or processing artifacts (1.5-12.5%). Corrective measures, including rescanning, xylene- or alcohol-based coverslip cleaning, and recoverslipping, resolved >95% of cases with cover-slip or mounting media issues, with limited tissue loss. Final slides were stored as pyramidized OME-TIFF files at 0.248 µm/pixel, averaging 2 GB per slide. In an offline Python pipeline, a Google Vision-derived OCR model yielded superior barcode read rate and accuracy on histology slides, outperforming Microsoft's OCR, Tesseract (v4+), and a Keras-based CRNN baseline. Conclusion: Our large-volume and high-throughput WSI archival workflow delivers a scalable imaging pipeline that digitized a 20-year academic tissue biobank and provides on-demand access to AI-ready, high-quality quality full-resolution slide images. This digital image repository will support large-scale digital pathology research to discover biomarkers in genitourinary cancer.
利益披露 Disclosure
F. Kabir, None.. M. Shavakhi, None.. A. Parvatikar, None.. A. Cesaire, None.. E. Phillips, None.. R. Trowbridge, None.. L. Ascione, None.. P. Barrios, None.. M. Eid, None.. J. Lee, None.. S. Signoretti, None.. E. V. Allen, None.. A. Choudhury, None.. L. Hoang, None.. T. K. Choueiri, None.. J. Wala, None.

← 返回 AACR 2026 检索