Advancing Vision Without Human Supervision Through Multimodal and Unsupervised Learning

DSpace Repositorium (Manakin basiert)


Dateien:

Zitierfähiger Link (URI): http://hdl.handle.net/10900/181694
http://nbn-resolving.org/urn:nbn:de:bsz:21-dspace-1816941
http://dx.doi.org/10.15496/publikation-123016
Dokumentart: Dissertation
Erscheinungsdatum: 2026-07-20
Sprache: Englisch
Fakultät: 7 Mathematisch-Naturwissenschaftliche Fakultät
Fachbereich: Informatik
Gutachter: Kuehne, Hilde (Prof. Dr.)
Tag der mündl. Prüfung: 2026-05-08
DDC-Klassifikation: 004 - Informatik
Schlagworte: Maschinelles Sehen , Unüberwachtes Lernen
Freie Schlagwörter: Selbstüberwachtes Lernen
Videoverstehen
Vision-Language-Modelle
Multimodales Lernen
Multimodal Learning
Unsupervised Learning
Self-supervised Learning
Computer Vision
Vision-language Models
Video Understanding
Lizenz: http://tobias-lib.uni-tuebingen.de/doku/lic_ohne_pod.php?la=de http://tobias-lib.uni-tuebingen.de/doku/lic_ohne_pod.php?la=en
Zur Langanzeige

Inhaltszusammenfassung:

Der Aufbau robuster, generalisierbarer Modelle, die in der Praxis über viele Anwendungen hinweg funktionieren, erfordert Lernverfahren, die nicht auf manuelle Beschriftungen angewiesen sind. Menschliche Annotationen sind kostspielig und zeitaufwendig, können Annotationsfehler beinhalten und können weder in großem Maßstab erstellt werden noch mit einer sich ständig wandelnden Welt Schritt halten. Gleichzeitig stehen im Internet große Mengen unbeschrifteter Daten über viele Modalitäten hinweg (z. B. Video, Audio, Text) zur Verfügung. Diese Dissertation verfolgt einen label-freien Ansatz für das Lernen von Bild- und Videorepräsentationen, der auf den Konzepten des multimodalen und des unüberwachten Lernens aufbaut. Diese Arbeit betrachtet dabei verschiedene Phasen des multimodal, unüberwachten Lernens: von der Erstellung hochwertiger Datensätze für das Vortraining über das Training grosser, allgemein einsetzbarer Modelle, deren Anpassung an spezifische Aufgaben und Domänen, bis hin zur Verbesserung von Evaluationsbenchmarks. Sie basiert auf drei komplementären Ideen: (1) der Nutzung eines multimodalen Vortrainings, in dem fusionierte multimodale Repräsentationen gelernt werden, und der Nutzung multimodaler Modelle als Pseudo-Supervisoren für zur Anpassung an bestimmte Aufgaben oder Domänen; (2) der Nutzung von Sprache als leistungsfähigem Repräsentationsraum, indem visuelle Signale in Textbeschreibungen übersetzt und die Schlussfolgerungsfähigkeiten großer Sprachmodelle ausgenutzt werden; und (3) der Anpassung von Loss-Funktionen, Architekturen und Pipelines für unüberwachtes Lernen in verschiedenen Phasen. Konkret führen wir einen Sorting-Loss für das selbstüberwachte Training ein, welcher die lokale Nachbarschaftsstruktur eines Repräsentationsraums im Vergleich zum häufig verwendeten Contrastive Loss verbessert. Anschließend konstruieren wir multimodale Video-Audio-Text-Repräsentationen, die flexibel jede Teilmenge von Modalitäten in einer gemeinsamen Repräsentationen fusionieren können. Daraufhin wird eine Pipeline vorgestellt, die Vision-Language-Modelle für das Videoverstehen – insbesondere für die Bewegungsklassifikation – anpasst, indem ausschließlich unbeschriftete Videos zusammen mit entsprechenden Labeln, aber ohne direkte Zuordnung, verwendet werden, wodurch der Bedarf an paarweiser Annotation entfällt. Des weiteren wurde eine Pipeline entwickelt, die Vision-Language-Modelle an spezifische Domänen (z. B. Kochen) anpasst, wobei nur domänenspezifische Textdaten genutzt werden – ohne jegliche Zieldomänen-Bilddaten oder gepaarte Annotationen. Um die Qualität von Video-Audio-Text-Daten im Web-Maßstab für das unüberwachtes, multimodale Training zu erhöhen, wurde eine Methode entwickelt, mit der verrauschte Untertitel von webbasierten Videos in hochwertige Beschreibungen umgewandelt werden, indem ein großes Sprachmodell das Video ausschließlich im Sprachraum interpretiert. Mit derselben Idee der sprachbasierten Verarbeitung wurde eine Methode zur Diagnose von visuelles Biases – etwa Objekt- oder Einzelframe-Bias – in gängigen Video-Evaluationsbenchmarks entwickelt, um eine robustere Bewertung aktueller Videomodelle zu ermöglichen. Schliesslich wird eine Architektur vorgestellt, die räumlich-zeitliche Aktivitätslokalisierung unter Verwendung ausschließlich webbasierter multimodaler Video-Text-Daten und ohne jegliche lokala Annotation lernt. Diese wird schliesslich ergänzt durch eine Methode, die Aktivitätslokalisierung auch durch vortrainierte multimodale Modelle ohne weiteres Training lediglich durch architektonische Modifikationen erlaubt. Insgesamt ermöglichen die genannten Beiträge das Training von Modellen zum Bild- und Videoverstehen ohne die Notwendigkeit manueller Beschriftungen zu verbessern. Die entwickelten Methoden für label-freies Training zeigen dabei, dass es möglich ist, gute Repräsentationsfähigkeiten ohne menschliche Annotationen zu lernen und stellen Werkzeuge vor, die die Qualität der Trainingsdaten verbessern und verlässlichere Evaluationsbenchmarks schaffen.

Abstract:

Building robust, generalizable models that work in the wild across many applications requires learning that does not rely on manual labels. Human annotations are costly and time-consuming. Labels inherit human biases and conventions, reflect taxonomic choices and annotation artifacts, and cannot be obtained at scale or keep pace with an always-changing, long-tailed world. Meanwhile, vast unlabeled data across many modalities (e.g., video, audio, text) are freely available online. Humans naturally learn without explicit teaching by discovering regularities in unimodal and multimodal inputs, integrating information across modalities, and using language to reason about them. This thesis advances a label-free path for image and video understanding, building on ideas of multimodal learning, reasoning in language space, and unsupervised learning. This thesis contributes to different main stages of building vision models without labels: from creating high-quality datasets for pretraining, to training strong general-purpose models, adapting them to specific tasks and domains, exploring models' dense properties, and finally improving evaluation benchmarks -- all without any human annotations. We develop three complementary ideas: (1) leverage multimodal pretraining to go beyond simple modality-alignment by learning fused multimodal representations, enabling dense spatio-temporal grounding capabilities, and using multimodal models as pseudo-supervisors for task/domain adaptation; (2) use language as the powerful operating space by translating visual signals into text descriptions and exploiting the strong reasoning abilities of large language models; and (3) design losses, architectures, and pipelines for unsupervised learning across different stages, such as pretraining, domain/task adaptation, and spatio-temporal grounding. Concretely, we introduce an ordering-based contrastive objective for self-supervised training that improves the local neighborhood structure of the learned embedding space compared to a commonly used contrastive loss. We then construct truly multimodal video–audio–text representations that can flexibly fuse any subset of modalities into a single joint embedding. Then, we present a pipeline that adapts image–language models to activity-centric video understanding, namely action recognition, using only unlabeled videos together with action dictionaries, eliminating the need for paired supervision. We also develop a pipeline that adapts vision–language models to specific domains (e.g., cooking) using only domain text data, without even target-domain vision data or any paired annotations. To increase the quality of of web-scale video–audio–text data for multimodal pretraining without human curation, we propose a method to convert noisy subtitles of web instructional videos into dense, high-quality, human-like descriptions by using LLM to reason about the videos entirely in language space. Using the same language-space reasoning idea, we propose a method to diagnose and mitigate representation biases, such as object or single-frame bias, in popular video evaluation benchmarks, enabling a more robust assessment of video understanding capabilities. We then propose an architecture that learns spatio-temporal activity localization without any grounding supervision, using only web multimodal video–text data. Finally, we unlock activity-grounding abilities in pretrained multimodal models entirely without training, via lightweight architectural modifications. Together, these contributions advance image and video understanding without reliance on manual labels. We develop methods for label-free model pretraining and adaptation, reveal dense grounding capabilities without human annotations, and introduce unsupervised tools that improve pretraining data quality and create fairer, more reliable evaluation benchmarks.

Das Dokument erscheint in: