PHYSICS-BASED ADAPTATION OF FOUNDATION MODELS FOR APPLICATIONS IN OPTICAL COHERENCE TOMOGRAPHY
Lev Matveev1, Denis Nikoshin1,2, Daniil Mikhailenko2, Nataliya Matveeva2,1, Konstantin Yashin3, Elena Kiseleva3, Sovetsky Alexander1, Radik Zinatullin3, Olga Streltsova3, Petrova Ksenia4, Mariam Novgorodskaya3, Maria Kulagina1,4, Kulagin Pavel1,4, Artem Grishin3, Liudmila Kukhnina3, Svetlana Korikova3, Maxim Ryabkov3, Vladimir Y. Zaitsev1, Alexander Matveyev1; 1Institute of Applied Sciences RAS, Nizhny Novgorod, Russia; 2HSE University, Nizhny Novgorod, Russia; 3Privolzhsky Research Medical University, Nizhny Novgorod, Russia; 4N.I. Lobachevsky Research State University, Nizhny Novgorod, Russia
Abstract
Recently, various types of foundation models have experienced rapid development. A defining feature of these models is that they are trained on broad, universal datasets and do not always require weight adjustments through fine-tuning; instead, their integration heavily relies on the principles of in-context learning. These models can be broadly divided into two classes: ViT-based computer vision foundation models (such as the Segment Anything Model, or SAM, family) and Vision-Language Models (VLMs) that process both text and images (e.g., Gemini, Gemma). Furthermore, VLMs can be categorized into open-source models that can be executed locally (like Gemma and MedGemma) and proprietary models accessible only via API (such as the Gemini family).
The primary limitation of all these models is their intrinsic lack of "physics-based" reasoning. They perceive scenes merely as standard visual pictures rather than as physical data, and none inherently possess the capability to process physically formed images. While open-source models like MedSAM offer pathways to incorporate this physical knowledge through fine-tuning, proprietary API-based models do not allow for weight updates outside their context window.
However, we demonstrate that these limitations can be successfully circumvented. Foundation models fit seamlessly into broader computational pipelines when utilized in an in-context learning mode as modular processing blocks. Specifically, they can be integrated into physics-informed pipelines at various stages, acting as both preprocessing and postprocessing units.
In this lecture, we will demonstrate a comprehensive range of pipeline variants that enable the physics-based adaptation of different foundation models across various optical coherence tomography (OCT) tasks. We will show how ViT models from the SAM family can be utilized for the preprocessing of OCT scans, enabling robust feature extraction and the subsequent derivation of segment parameters. Additionally, we will illustrate how VLMs, using the Gemini family as an example, can be integrated for the postprocessing of physics-based processed maps by assessing them with in-context knowledge prompts, including qualitative clinical rules. Our specific applications will highlight the construction of these pipelines for the diagnostic assessment of skin, urethra, and brain OCT scans.
In conclusion, we will demonstrate how multi-agent systems built upon these foundation models bring us closer to solving complex inverse-physics problems. Specifically, we will show how these systems can reconstruct the spatial distributions of optical scatterers from OCT scans to create digital phantoms. This advancement paves the way toward the realization of non-invasive virtual histology based on OCT.
Acknowledgment: This work was supported by the RSCF grant № 25-12-20032 "New Approaches to the Development of Algorithms for Analyzing OCT Scans: Modification and Optimization of Large Models Based on Physical Principles and Conditions of OCT Signal Formation"
Speaker
Lev A. Matveev
Institute of Applied Sciences, Russia Academy of Sciences
Russia
Discussion
Ask question