Preparing for LSST: An Autoencoder Approach to Anomalous AGN Variability
TALES doctoral candidate Natale De Bonis has recently published a scientific paper presenting a Machine Learning (ML) method for automatically identifying unusual Active Galactic Nuclei (AGN) from their variability.
The paper was submitted to Astronomy & Astrophysics (A&A) at the end of March and was accepted for publication at the end of June.
The world of astronomy is changing very rapidly; with the Vera C. Rubin Observatory’s Legacy Survey of Space and Time (LSST), we will have a dynamic view of the universe never seen before. With such a vast amount of data, identifying the relatively small number of sources displaying unusual behaviour will be an important challenge, since these objects may reveal new or unexpected physical phenomena.
Finding these sources is not an easy task, since astronomical light curves are affected by irregular cadences, different durations, and systematic effects.
To address these problems, we developed an unsupervised Machine Learning method based on an Autoencoder. Rather than analysing the light curves directly, we first describe each source through a set of statistical and physical features that capture different aspects of its variability. The Autoencoder then learns the typical properties of the AGN population and attempts to reproduce them. Sources whose properties differ substantially from those learned by the model are reconstructed less accurately. These sources can then be flagged as potential anomalies for further investigation.
We applied this approach to a large sample of AGN observed in the COSMOS field and focused our attention on obscured AGN, which tend to be harder to characterize from their optical variability alone, since their central regions are hidden from direct view by surrounding dust.
An important part of the study was understanding why the Autoencoder considers a source anomalous. By analysing the contribution of the different input features, we identified a smaller and more compact set of features that retains a performance comparable to that of the full feature set. This makes the method more efficient while also providing insight into which properties of the variability are responsible for the anomalous classification.
These results demonstrate the potential of Machine Learning to efficiently explore large astronomical datasets and identify sources that deserve further investigation. This will become increasingly important for next-generation surveys such as LSST, where the enormous volume of time-domain data will make automated methods essential for selecting scientifically interesting objects for further study and follow-up observations.

Figure Caption: Visual representation of the obscured AGN population based on their light-curve properties. Points are color-coded according to their reconstruction error (RE) score, while black circles highlight sources flagged as anomalous by the autoencoder.



