Deep learning methods for satellite image timeseries
File(s)
Author(s)
Tarasiou, Michail
Type
Thesis
Abstract
The practice of Earth Observation (EO) plays a critical role in monitoring human activities and their environmental impacts through remote sensing technologies. This is crucial for designing interventions that enhance societal prosperity and resilience. Recent advancements in Deep Neural Networks (DNNs), coupled with increasingly available EO data, present a unique opportunity for improving future production systems. However, current research crucially overlooks unique attributes of satellite imagery and suffers from a lack of publicly accessible, large-scale annotated datasets. This PhD thesis explores the potential of automating EO tasks within a DNN framework. To this end, several key aspects of the problem are explored.
Firstly, a dataset generation tool is introduced that standardizes and simplifies the processing of Sentinel-2 data for DNN applications. DeepSatData is designed to be lightweight and can be adapted to meet various research needs through the Python library ecosystem.
Furthermore, the Temporo-Spatial Vision Transformer (TSViT), the first fully-attentional model optimized for processing Satellite Image Time Series (SITS), is introduced. TSViT employs input factorization and specific inductive biases that improve its discriminative capability, substantially outperforming existing models in large-scale land cover recognition benchmarks.
To take advantage of large scale SITS data readily available, EmbeddingEarth, a self-supervised pre-training method, is developed to boost the performance of land cover classification at the pixel level. This method allows models, pre-trained in different regions, to be effectively fine-tuned on new areas, indicating the potential for developing foundational models for EO.
Lastly, a fully supervised pre-training method is introduced, targeting performance improvements at object boundaries for semantic segmentation. Context-Self Contrastive Pre-training, using no additional data, enhances the performance of all state-of-the-art networks on land cover recognition, particularly at fine-grain boundaries.
Firstly, a dataset generation tool is introduced that standardizes and simplifies the processing of Sentinel-2 data for DNN applications. DeepSatData is designed to be lightweight and can be adapted to meet various research needs through the Python library ecosystem.
Furthermore, the Temporo-Spatial Vision Transformer (TSViT), the first fully-attentional model optimized for processing Satellite Image Time Series (SITS), is introduced. TSViT employs input factorization and specific inductive biases that improve its discriminative capability, substantially outperforming existing models in large-scale land cover recognition benchmarks.
To take advantage of large scale SITS data readily available, EmbeddingEarth, a self-supervised pre-training method, is developed to boost the performance of land cover classification at the pixel level. This method allows models, pre-trained in different regions, to be effectively fine-tuned on new areas, indicating the potential for developing foundational models for EO.
Lastly, a fully supervised pre-training method is introduced, targeting performance improvements at object boundaries for semantic segmentation. Context-Self Contrastive Pre-training, using no additional data, enhances the performance of all state-of-the-art networks on land cover recognition, particularly at fine-grain boundaries.
Version
Open Access
Date Issued
2023-08
Date Awarded
2024-05
Copyright Statement
Creative Commons Attribution Licence
License URL
Advisor
Zafeiriou, Stefanos
Publisher Department
Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)