USCMS Researcher: Matteo Marchegiani
Postdoc dates: Jul 2025 - Jun 2027
Home Institution: Carnegie Mellon University
Project: GNN-based End-to-End Reconstruction in the CMS Phase 2 High-Granularity Calorimeter
2026-2027 Transformer-based End-to-End Reconstruction in the CMS Phase 2 High-Granularity Calorimeter
Building on the GNN-based reconstruction developed in the first year, the goal of this second project is to study transformer-based architectures, and specifically MaskFormers, as an alternative approach to end-to-end reconstruction in the HGCAL. A transformer with full self-attention is equivalent to a fully-connected message-passing GNN, offering greater expressivity than the local graph connectivity used by GNNs, and masked attention mitigates the quadratic complexity that would otherwise result from attending over the up to 200k hits expected per event at 200 pile-up. The project starts by interfacing the existing HGCAL training datasets with a MaskFormer model, aggregating information from multiple detector hits into composed objects, the topoclusters, to keep the self-attention matrix computationally tractable. The model is then trained with progressively higher pile-up (0, 30, and 200 collisions), monitoring memory and compute costs against the CMS Phase-2 offline processing budget, before its energy, position and time resolution and response are benchmarked against the GNN baseline using the same performance-metric infrastructure developed in the prior award.
2026 Project proposal
2025-2026 GNN-based End-to-End Reconstruction in the CMS Phase 2 High-Granularity Calorimeter
The High-Luminosity LHC (HL-LHC) will deliver up to 200 simultaneous interactions per bunch crossing (pile-up), posing an exceptional challenge for particle-shower reconstruction in the CMS Phase 2 High-Granularity Calorimeter (HGCAL). The goal of this project is to develop Graph Neural Network (GNN) based algorithms for fast and efficient end-to-end reconstruction of particle showers in the HGCAL, capable of meeting this challenge. Building on an existing proof-of-concept model that uses GravConv layers together with the Object Condensation loss function to reconstruct energy clusters from simulated di-tau decays with zero pile-up, the project extends this approach to realistic HL-LHC pile-up conditions. This requires deriving a scalable, pile-up-aware training dataset by combining full-simulation truth information from FineCalo with a dedicated library of minimum-bias events and a merging algorithm that reconciles truth information between the hard-scatter and pile-up interactions, as well as re-engineering and optimizing the Object Condensation loss function to overcome GPU memory limitations at high pile-up and to improve physics performance, exploring alternatives such as the modified differential multiplier method and the influencer loss.
2025 Project proposal
More information: My project proposal
Mentors:
-
Matteo Cremonesi - (Carnegie Mellon University)
- 25 May 2026 - "GNN-based end-to-end reconstruction in the CMS Phase-2 High-Granularity Calorimeter", Matteo Marchegiani, CHEP 2026
- 15 May 2026 - "GNN-based end-to-end reconstruction in the CMS Phase-2 High-Granularity Calorimeter", Matteo Marchegiani, DPG Plot Approval Meeting
- 7 May 2026 - "GNN-based end-to-end reconstruction in the CMS Phase-2 High-Granularity Calorimeter", Matteo Marchegiani, TICL Reconstruction Working Meeting
- 23 Apr 2026 - "HGCAL reconstruction with Graph Neural Networks", Matteo Marchegiani, TICL Reconstruction Working Meeting
- 26 Mar 2026 - "HGCAL reconstruction with Graph Neural Networks", Matteo Marchegiani, TICL Reconstruction Working Meeting
- 24 Mar 2026 - "HGCAL simulation with pileup and merging algorithm", Matteo Marchegiani, ML4RECO
- 10 Feb 2026 - "Maskformers for offline reconstruction", Matteo Marchegiani, ML4RECO
- 3 Dec 2025 - "Studies on time resolution in GNN-based reco", Matteo Marchegiani, HGCAL DPG
- 2 Dec 2025 - "Studies on time resolution with the latest HGCAL GNN model", Matteo Marchegiani, ML4RECO
- 18 Nov 2025 - "Update on performance of the latest HGCAL GNN model", Matteo Marchegiani, ML4RECO
- 13 Oct 2025 - "ML4RECO: GNN and Transformer Based HGCAL Reconstruction", Matteo Marchegiani, CMS Machine Learning Town Hall
- 17 Sep 2025 - "Single-particle energy resolution with latest GNN model", Matteo Marchegiani, ML4RECO
- 27 Aug 2025 - "Updated GNN training after bugfixes in CMSSW 15_0_X", Matteo Marchegiani, ML4RECO
- 16 Jul 2025 - "FineCalo SimTrack Reconstruction Error", Matteo Marchegiani, ML4RECO
Current Status
2026 Q2
- Working on training of GNN model with 200 pileup simulation
- First training with 30 PU using SimCluster features as input
- Study new alternative ML architectures for HGCAL reconstruction
- Generated large 0 PU dataset with RecHits and TICL objects to train a large model
- New training using recHits features as input on 1000 events: incidence matrix regression and regression of cluster properties
- Optimization of the model’s architecture and parameters
- Publication of DP Note on GNN-based HGCAL reconstruction
- GravNet model trained with object condensation loss with 0 PU simulation
- MC truth: CMSSW-native SimClusters
- Presentation at CHEP 2026
2026 Q1
- Progress on optimized GNN model
- Study impact of pileup on the training of the GNN-based reconstruction algorithm
- GPU Memory profiling of GNN model trained on dataset with 5x more hits
- Integration of CUDA kernels from FastGraphCompute to speed up KNN and object condensation loss
- Working on training of GNN model with 200 pileup simulation
- First HGCAL simulation with FineCalo + merging algorithm in CMSSW_15_1_0
- Simulated 10 single-electron events + 30 PU interactions from minimum bias events
- Implementation of custom NANO step to save true clusters containing SimHits from primary interaction and pileup
- At 30 PU, the average number of RecHits is ~60k, with ~8k true clusters
- Study a dedicated implementation of the merging algorithm which is compatible with pileup
- Study a dedicated pileup mixing library to save the event history for minimum bias events in the merging
- Study new alternative ML architectures for HGCAL reconstruction
- Masked transformers to reduce the quadratic complexity of self-attention
- Generated 0 PU dataset with RecHits and TICL objects: 2D LayerClusters, 3D Tracksters and TICL Candidates
- Computing the incidence matrix between RecHits and true particles using sparse tensor representation
- First proof-of-concept training using recHits features as input: incidence matrix regression and regression of cluster properties
2025 Q4
- Progress
- Study energy, position and time resolution of the GNN clustering using simulated photons, pions and tau leptons
- Study issue in time resolution in CMSSW_15_X: change in the timing simulation with respect to CMSSW_11_X
- Finalize publication on performance metrics of GNN clustering with 0 pileup simulation
- Study the impact of increased pileup on the GNN training
- Memory profiling of the training with a dataset including ~250k reconstructed hits
- Working on pileup simulation: generate hard process alongside minimum bias events
2025 Q3
- Progress
- Learned how to train the GNN employing GravConv layers in combination with the object condensation loss for the reconstruction of energy clusters in the HGCAL
- Studied the performance and energy resolution of the GNN-based reconstruction in zero-pileup environments, considering simulated photons, pions and tau leptons
- Ported the training datasets production to CMSSW_15_1_X, making use of the FineCalo simulation
- Ported the GNN training to Pytorch 2.6.0 and CUDA 12.4
- Study memory profiling of the model on GPU
Contact me: