![]() |
Abstract EANA2026-245 |
|
Exploring Technosignatures Using Pyrolysis-Gas Chromatography Mass Spectrometry and Machine Learning
The search for life in space is a central goal of astrobiology, and a large part of current astrobiological research aims to maximize our ability to accurately infer the provenance of samples analyzed in planetary environments. One such technique uses pyrolysis-gas chromatography coupled to mass spectrometry (py-GC-MS) combined with machine learning (ML) algorithms. Py-GC-MS provides a detailed inventory of the molecular content of samples. ML then analyzes the molecular patterns to learn how to distinguish samples of different classes. In several recent studies this technique has been able to determine the biogenicity of samples to over 90% accuracy.
Until now, using ML to classify samples has focused on samples that fall into two broad categories, biotic vs. abiotic, but there is a gap in understanding how biotic matter that has been heavily processed by human technological intervention will be classified. Our research advances this biosignature detection method by expanding the sample set to incorporate ultra-processed products—in other words, technosignatures—into the overall project, prioritizing human-made goods acquired from typical supermarkets and stores. Our 73 samples included a variety of processed items, from chips, pastas, and breads to fresh fruit and vegetable produce; we also expanded into non-vegan cosmetic items, such as retinols, make-up items, and hair products.
Preliminary results: Unsupervised clustering shows that ultra-processed samples tend to cluster separately from biotic and abiotic samples. ML classifiers trained on classic biotic and abiotic specimens tend to categorize most ultra-processed samples as biotic, rather than abiotic. In general, less-processed samples (e.g., pastas and breads) had a high average classification probability (>0.70). On the other hand, the more processed samples (e.g., snack foods and cereals) tended to have a lower average classification probability (0.5–0.7), meaning the ML algorithm is less able to determine the biogenicity of heavily processed samples.