Exploring unsupervised top tagging using Bayesian inference

Alvarez, Ezequiel; Szewc, Manuel; Szynkman, Alejandro; A. Tanco, Santiago; Tarutina, Tatiana

doi:10.21468/SciPostPhysCore.6.2.046

SciPost Physics Core

Exploring unsupervised top tagging using Bayesian inference

Ezequiel Alvarez, Manuel Szewc, Alejandro Szynkman, Santiago A. Tanco, Tatiana Tarutina

SciPost Phys. Core 6, 046 (2023) · published 28 June 2023

doi: 10.21468/SciPostPhysCore.6.2.046
pdf
Submissions/Reports

Abstract

Recognizing hadronically decaying top-quark jets in a sample of jets, or even its total fraction in the sample, is an important step in many LHC searches for Standard Model and Beyond Standard Model physics as well. Although there exists outstanding top-tagger algorithms, their construction and their expected performance rely on Montecarlo simulations, which may induce potential biases. For these reasons we develop two simple unsupervised top-tagger algorithms based on performing Bayesian inference on a mixture model. In one of them we use as the observed variable a new geometrically-based observable $\tilde{A}_{3}$, and in the other we consider the more traditional $\tau_{3}/\tau_{2}$ $N$-subjettiness ratio, which yields a better performance. As expected, we find that the unsupervised tagger performance is below existing supervised taggers, reaching expected Area Under Curve AUC $\sim 0.80-0.81$ and accuracies of about 69\% $-$ 75\% in a full range of sample purity. However, these performances are more robust to possible biases in the Montecarlo that their supervised counterparts. Our findings are a step towards exploring and considering simpler and unbiased taggers.

TY  - JOUR
PB  - SciPost Foundation
DO  - 10.21468/SciPostPhysCore.6.2.046
TI  - Exploring unsupervised top tagging using Bayesian inference
PY  - 2023/06/28
UR  - https://scipost.org/SciPostPhysCore.6.2.046
JF  - SciPost Physics Core
JA  - SciPost Phys. Core
VL  - 6
IS  - 2
SP  - 046
A1  - Alvarez, Ezequiel
AU  - Szewc, Manuel
AU  - Szynkman, Alejandro
AU  - A. Tanco, Santiago
AU  - Tarutina, Tatiana
AB  - Recognizing hadronically decaying top-quark jets in a sample of jets, or even its total fraction in the sample, is an important step in many LHC searches for Standard Model and Beyond Standard Model physics as well. Although there exists outstanding top-tagger algorithms, their construction and their expected performance rely on Montecarlo simulations, which may induce potential biases. For these reasons we develop two simple unsupervised top-tagger algorithms based on performing Bayesian inference on a mixture model. In one of them we use as the observed variable a new geometrically-based observable $\tilde{A}_{3}$, and in the other we consider the more traditional $\tau_{3}/\tau_{2}$ $N$-subjettiness ratio, which yields a better performance. As expected, we find that the unsupervised tagger performance is below existing supervised taggers, reaching expected Area Under Curve AUC $\sim 0.80-0.81$ and accuracies of about 69\% $-$ 75\% in a full range of sample purity. However, these performances are more robust to possible biases in the Montecarlo that their supervised counterparts. Our findings are a step towards exploring and considering simpler and unbiased taggers.
ER  -

@Article{10.21468/SciPostPhysCore.6.2.046,
	title={{Exploring unsupervised top tagging using Bayesian inference}},
	author={Ezequiel Alvarez and Manuel Szewc and Alejandro Szynkman and Santiago A. Tanco and Tatiana Tarutina},
	journal={SciPost Phys. Core},
	volume={6},
	pages={046},
	year={2023},
	publisher={SciPost},
	doi={10.21468/SciPostPhysCore.6.2.046},
	url={https://scipost.org/10.21468/SciPostPhysCore.6.2.046},
}

Cited by 2

Ontology / Topics

See full Ontology or Topics database.

Jets Machine learning (ML) Top quark

Authors / Affiliations: mappings to Contributors and Organizations

See all Organizations.

¹ Ezequiel Alvarez,
² Manuel Szewc,
³ Alejandro Szynkman,
³ Santiago A. Tanco,
¹ Tatiana Tarutina

Funders for the research work leading to this publication