FILTR: Extracting Topological Features from Pretrained 3D Models

LIX, Ecole Polytechnique, IP Paris
CVPR 2026 ยท Oral
TL;DR

We present the first study of topological understanding in pretrained 3D point cloud encoders, introduce DONUT, a benchmark of 30K topologically annotated shapes, and propose a framework to extract the topology of a point cloud from pretrained encoder representations.

Abstract

Recent advances in pretraining 3D point cloud encoders (e.g., Point-BERT, Point-MAE) have produced powerful models, whose abilities are typically evaluated on geometric or semantic tasks. At the same time, topological descriptors have been shown to provide informative summaries of a shape's multiscale structure. In this paper we pose the question whether topological information can be derived from features produced by 3D encoders. To address this question, we first introduce DONUT, a synthetic benchmark with controlled topological complexity, and propose FILTR (Filtration Transformer), a learnable framework to predict persistence diagrams directly from frozen encoders. FILTR adapts a transformer decoder to treat diagram generation as a set prediction task. Our analysis on DONUT reveals that existing encoders retain only limited global topological signals, yet FILTR successfully leverages information produced by these encoders to approximate persistence diagrams. Our approach enables, for the first time, data-driven extraction of persistence diagrams from raw point clouds through an efficient learnable feed-forward mechanism.

Method

We evaluate the topological understanding of pretrained 3D point-cloud encoders through three complementary tasks: predicting the number of connected components of the underlying shape, predicting its genus, and measuring how well encoder features align with descriptors derived from persistence diagrams.

Tasks 1 & 2

Probing global topology

We probe frozen encoder features to predict the number of connected components of the underlying shape, then its genus, which is a more demanding test of global topological structure. Each transformer block is probed separately with linear layers on point clouds from DONUT.

Input point cloudPoint cloud
Transformer encoderself-supervised reconstruction
Block 1
Block 2
Block 3
Reconstructed point cloudReconstruction
Underlying meshUnderlying topology
  1. 1Input: a point cloud sampled from a 3D shape.
  2. 2A pretrained transformer encoder, trained with self-supervised objectives such as masked reconstruction, processes it block by block.
  3. 3The encoder stays frozen. Linear probes on each block's features predict the number of connected components and the genus of the underlying shape.

Finding. Encoders have a limited understanding of the global structure of point clouds: probing accuracies are only marginally better than a PointNet trained from scratch.

Background

Persistent homology, a multiscale descriptor

Probing only tests global labels. To look at topology across scales, we use persistent homology: it tracks when topological features such as connected components and loops appear and disappear along a filtration, and summarizes them in a persistence diagram.

Point cloud Persistence barcode H0 H1 Persistence diagram Birth Death Loop Connected components more salient structure Filtration Process
  1. 1Start from a point cloud: every point is its own connected component.
  2. 2Grow a ball around each point. When balls overlap, points connect, components merge, and their H0 bars end.
  3. 3As the radius grows, more components merge.
  4. 4All points are now connected: a single component remains.
  5. 5The edges close a cycle: a loop is born, starting an H1 bar.
  6. 6Triangles start filling the loop in.
  7. 7Once the loop is completely filled, it dies and its H1 bar ends.
  8. 8Each bar becomes a (birth, death) point of the persistence diagram. The farther from the diagonal, the more salient the feature.
Task 3

Feature space similarity

We measure the alignment between encoder features and vectorized persistence diagrams with Centered Kernel Alignment (CKA). This parameter-free comparison assesses whether multiscale topological information is present in the learned representations, and how accessible it is.

Point cloudPoint cloud
Persistence diagram
Encode
Vectorize
Encoder feature manifold Persistence feature manifold
Centered kernel Alignment
Aligned encoder feature manifold Aligned persistence feature manifold
  1. 1Two views of the same shape: its point cloud and its persistence diagram.
  2. 2The point cloud is encoded by the frozen 3D encoder; the diagram is turned into a vector.
  3. 3Across the dataset, each side defines a feature space.
  4. 4CKA measures how well the two feature spaces align, without training anything.

Finding. Encoder features correlate with vectorized persistence diagrams: 3D models might implicitly capture local, multiscale topological structure.

DONUT

Dataset Of maNifold strUcTures

30Kmeshes
0โ€“10genus
1โ€“6connected components
DONUT dataset samples

DONUT is a synthetic benchmark of 30K meshes annotated with two key topological quantities: the total genus of each sample and its number of connected components. The benchmark spans genera from 0 to 10 and between 1 and 6 connected components, while keeping the distributions balanced. Our probing experiments fully rely on this benchmark to evaluate what pretrained 3D point cloud encoders retain about topology, and we also use DONUT to train the encoder used in FILTR.

FILTR

Filtration Transformer

FILTR architecture

FILTR is a DETR-inspired framework for predicting persistence diagrams directly from point clouds in a feed-forward manner. A frozen pretrained 3D encoder first produces point-cloud features and associated 3D positional information, which are projected to the decoder dimension and used to condition a transformer decoder through cross-attention. The decoder processes a fixed set of learned queries and outputs unordered persistence pairs together with existence scores, allowing the model to represent both genuine topological features and empty slots.

FILTR enforces the birth-death ordering in its pair parameterization and is trained with a set-prediction objective combining Hungarian matching, regression on matched pairs, an existence loss, and a diagonal regularizer for unmatched predictions. This design lets FILTR leverage frozen 3D features to approximate persistence diagrams efficiently, without iterative post-processing or end-to-end retraining of the backbone.

โˆ’73%prediction error vs. end-to-end counterparts, in the low-data regime
5.7Mtrainable parameters with a pretrained encoder, vs. ~8M for end-to-end baselines
FILTR persistence diagram generation results
Persistence diagrams generated by FILTR. Although FILTR is trained on DONUT, it still manages to accurately reconstruct persistence diagrams for more challenging point clouds without any fine-tuning.

BibTeX

@inproceedings{Martinez2026FILTR,
  title={FILTR: Extracting Topological Features from Pretrained 3D Models},
  author={Louis Martinez and Maks Ovsjanikov},
  booktitle={Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
  year={2026},
  url={https://arxiv.org/abs/2604.22334}
}