Online identity inference, impersonation and mitigation strategies
File(s)
Author(s)
Lepipas, Anastasios
Type
Thesis
Abstract
Online identities in social platforms and generative models are increasingly at risk as adversaries exploit unique identifiers for malicious ends. This thesis examines two complementary perspectives on such vulnerabilities: squatted usernames on social networks and attribute leakage in text-guided diffusion (T2I) models.
First, we systematically investigate how attackers capitalize on near-duplicate username registrations on X, confusing users and enabling large-scale impersonation. To illuminate this phenomenon, we develop the username-generation tool UsernameCrazy, inspecting hundreds of thousands of variants derived from celebrity accounts. Our measurements reveal a concerning ecosystem of suspended and active squatted usernames, many featuring profile pictures and names closely matching the originals to bolster deception. We further demonstrate how user mismentions and X’s search algorithm magnify these threats. To mitigate them, we propose SQUAD, a framework integrating UsernameCrazy, achieving a 94% F1-score in detecting suspicious squatted accounts.
While username squatting targets explicit digital identities, online personas are increasingly shaped by more subtle signals—particularly in the era of generative AI. We extend our analysis to emerging T2I models and their potential to leak sensitive traits-such as authorship or dementia status-through seemingly benign image outputs. By constructing adversarial pipelines that leverage image augmentation and text–image embedding models, we achieve up to 0.877% Top-5 accuracy in attributing images across 100 authors, and 0.75% accuracy in inferring dementia status (using the ADReSS dataset). These inferences are robust against diverse training sets, independent of classifier choice, and remain resilient to standard mitigation strategies.
Together, these findings reveal how adversaries can hijack identifiers-whether user-chosen names on a social platform or latent attributes embedded in generative images-to threaten security and privacy. By characterizing these overlooked risks and introducing detection and protection mechanisms, this thesis advances our collective understanding of modern identity vulnerabilities and highlights avenues to safeguard online communities.
First, we systematically investigate how attackers capitalize on near-duplicate username registrations on X, confusing users and enabling large-scale impersonation. To illuminate this phenomenon, we develop the username-generation tool UsernameCrazy, inspecting hundreds of thousands of variants derived from celebrity accounts. Our measurements reveal a concerning ecosystem of suspended and active squatted usernames, many featuring profile pictures and names closely matching the originals to bolster deception. We further demonstrate how user mismentions and X’s search algorithm magnify these threats. To mitigate them, we propose SQUAD, a framework integrating UsernameCrazy, achieving a 94% F1-score in detecting suspicious squatted accounts.
While username squatting targets explicit digital identities, online personas are increasingly shaped by more subtle signals—particularly in the era of generative AI. We extend our analysis to emerging T2I models and their potential to leak sensitive traits-such as authorship or dementia status-through seemingly benign image outputs. By constructing adversarial pipelines that leverage image augmentation and text–image embedding models, we achieve up to 0.877% Top-5 accuracy in attributing images across 100 authors, and 0.75% accuracy in inferring dementia status (using the ADReSS dataset). These inferences are robust against diverse training sets, independent of classifier choice, and remain resilient to standard mitigation strategies.
Together, these findings reveal how adversaries can hijack identifiers-whether user-chosen names on a social platform or latent attributes embedded in generative images-to threaten security and privacy. By characterizing these overlooked risks and introducing detection and protection mechanisms, this thesis advances our collective understanding of modern identity vulnerabilities and highlights avenues to safeguard online communities.
Version
Open Access
Date Issued
2025-04-06
Date Awarded
01/08/2025
License URL
Advisor
Demetriou, Soteris
Publisher Department
Department of Computing
Publisher Institution
Imperial College London
Qualification Level
Doctoral
Qualification Name
Doctor of Philosophy (PhD)
