|
Authors: E. Di Napoli, L. E. M. Zumbo, D. De Biase, G. Piegari, S. Papparella, V. Russo and O. Paciello
|
||||||
|
Résumé, analyse et commentaires |
||||||
|
Aucun.
|
||||||
|
Photo |
||||||
|
Aucune.
|
||||||
|
Analysis |
||||||
|
None.
|
||||||
|
Abstract |
Source |
|
|
|||
|
Canine cutaneous neoplasms are common and morphologically heterogeneous lesions whose diagnosis relies on integrating gross examination, cytology, and histopathology. This retrospective pilot study assessed the feasibility of a multimodal GPT-based large language model as an assistive, not autonomous, tool for standardized description, differential diagnosis generation, and classification support across this diagnostic workflow. Fifty-one histologically confirmed canine cutaneous tumors were retrospectively selected from the laboratory information system of the Veterinary Pathology Laboratory, University of Naples Federico II. For each case, de-identified gross photographs, digitized cytology, and representative histologic images were provided to the model using templated prompts. Model outputs were independently reviewed by two veterinary pathologists, who reached consensus on descriptive quality and diagnostic concordance with the histologic reference diagnosis. Final diagnostic outputs were classified as correct, partially correct, or incorrect. Strict accuracy was defined as the proportion of fully correct diagnoses, whereas broad accuracy combined correct and partially correct outputs considered diagnostically informative. Overall, the model achieved a strict diagnostic accuracy of 66.7% (34/51; 95% CI: 53.0-78.0) and a broad diagnostic accuracy of 90.2% (46/51; 95% CI: 79.0-95.7). Performance was highest in epithelial tumors and lower in mesenchymal and melanocytic tumors, in which the model more often identified broader diagnostic categories than specific histotypes. These findings suggest that GPT-based systems may support report standardization, descriptive consistency, and morphology-driven reasoning in veterinary pathology. However, reduced entity-level specificity, variable descriptive quality, and the risk of plausible but non-concordant outputs require strict human supervision and further validation before routine implementation.
|
||||||