CFM is a language-aligned concept foundation model that extracts spatially localized, visually grounded, and human-interpretable concepts at various granularities from images, organizes them into hierarchies, and automatically assigns names to them — enabling concept-based explanations for any downstream task that the foundation model can perform, such as classification, open-vocabulary segmentation, and captioning.
@article{wittenmayer2026cfm,
title = {CFM: Language-aligned Concept Foundation Model for Vision},
author = {Wittenmayer, Kai and Rao, Sukrut and Parchami-Araghi, Amin
and Schiele, Bernt and Fischer, Jonas},
journal = {arXiv preprint arXiv:2601.13798},
year = {2026}
}
Concepts and their names are extracted automatically from large-scale image data without manual curation. Some concepts or their automatically assigned names may be inaccurate or offensive; they do not reflect the views of the authors.
Matryoshka shells — shell 1 is the innermost (most general) concepts.
Node size = frequency · Edge length ∝ 1 / co-occurrence.