Computer Science

Permanent URI for this collectionhttps://hdl.handle.net/10315/38508

Browse

Recent Submissions

Now showing 1 - 20 of 144
  • Item type: Item , Access status: Open Access ,
    Investigating Illusory Motion Parallax Via Visual Cue Decoupling and Forward Models in VR
    (2026-07-24) Kordmodanlou, Daniel; Troje, Nikolaus
    Under natural viewing, stereopsis and motion parallax provide consistent information about three-dimensional structure, but stereoscopic displays can decouple them: observers moving before a stereoscopic display that lacks viewpoint-contingent updates perceive illusory parallactic motion. This thesis investigates that illusion in head-mounted virtual reality and tests whether it arises from predictive mechanisms that use stereoscopic depth to generate expectations about motion parallax. Across five two-interval forced-choice experiments, the illusion required binocular disparity, and its perceived direction aligned with negative parallax, consistent with a signed prediction error rather than an unsigned cue conflict. Its magnitude was modulated by scene context, downward viewing angle, and observer distance. These findings support a forward-model account: binocular disparity specifies depth and predicts head-coupled motion parallax, and the absent expected parallax yields a prediction error experienced as scene motion. The results inform how the visual system integrates conflicting depth cues and the design of stereoscopic displays.
  • Item type: Item , Access status: Open Access ,
    Towards Robust and Coherent Data Visualization Generation with Vision–Language Models
    (2026-07-24) Mahbub, Ridwan; Prince, Enamul Hoque
    Visual data stories have emerged as an effective medium for communicating data by integrating visualizations with coherent narratives. Data videos, a popular format of visual data stories, combine animated, chart-centric visualizations with synchronized narration. In this thesis, we introduce a novel task of data video generation for Vision Language Models (VLMs), along with a benchmark containing 328 real world data videos. We propose a multi-agent framework for generating better data videos, composed of planning and critic agents, designed to mimic the real-world process of data video generation. While these results highlight the potential of VLMs for generating coherent visual data stories, they also underscore the need to examine their robustness. Accordingly, we investigate how misleading visual designs as input can affect VLM behavior and performance, and ways to mitigate this effect. Together, these contributions not only extend the capabilities of VLMs into a new dimension of visual data storytelling, but also provide a systematic understanding of their robustness in general.
  • Item type: Item , Access status: Open Access ,
    Efficient Fine-Tuning of Foundation Models: Generalization, Disparity, and Subgroup Adaptation
    (2026-07-24) Parkhimchyk, Artur; Seyyed-Kalantari, Laleh
    Foundation models (FMs) perform well in vision but remain difficult to deploy when compute is limited and data are demographically imbalanced. This thesis makes two contributions. First, it studies demographic adaptation with parameter-efficient fine-tuning (PEFT) under realistic constraints where full fine-tuning is infeasible and available data are skewed. We evaluate five adaptation strategies, three pretraining paradigms, and multiple medical imaging datasets, covering 19 dataset-demographic variants, 490 configurations, and over 3,500 GPU hours. Results show that demographic adaptation can improve minority-group performance, especially with E2VPT, and clarify when PEFT offers strong efficiency--performance tradeoffs. Second, we propose Low-Rank Adaptation with Sinusoidal Projection Sampling (LoRA-SPS), a sinusoidal-projection low-rank adapter that samples from pretrained weights, decouples parameter count from embedding dimension, and improves projection rank stability while remaining mergeable for zero-overhead inference. Across vision and language tasks, LoRA-SPS matches or outperforms strong baselines using substantially fewer parameters.
  • Item type: Item , Access status: Open Access ,
    Avatar Fidelity and Partner Representation in Collaborative Extended Reality
    (2026-07-24) Ul Haq, Ehtisham; Allison, Robert S.
    Collaborative extended reality systems depend critically on how partners are represented. This thesis examines how avatar fidelity and partner representation shape coordination in shared immersive tasks. Two empirical studies compare representation conditions across asymmetric guidance and fast-paced shared-workspace interaction. The results show that more realistic or human-like representations do not uniformly improve collaboration. Instead, different representation conditions produce distinct coordination tradeoffs across task performance, errors, interference management, workload, social presence, and subjective experience. In gesture-dependent guidance, human-like appearance without corresponding behavioral support increases coordination costs. In fast-paced shared-workspace interaction, similar overall performance can emerge through different coordination strategies, including variations in selectivity, caution, and interference management. Together, the studies show that partner representation systematically shapes coordination processes and tradeoffs. It should therefore be understood as a design factor that determines what information collaborators can use and how they coordinate action, rather than simply as a means of increasing realism.
  • Item type: Item , Access status: Open Access ,
    A Semantic-Aware Reinforcement Learning Scheduler for Parameterized Deep Learning Jobs on Heterogeneous Multi-GPU Clusters
    (2026-07-24) Zhou, Zehao; An, Aijun
    Efficient scheduling of deep learning workloads in heterogeneous GPU clusters is increasingly challenging due to diverse workload characteristics and stringent memory constraints. Existing approaches often rely on coarse-grained metrics such as runtime or resource demand, which fail to capture workload semantics and lead to resource misallocation, where high-memory GPUs are occupied by low-demand jobs while memory-intensive workloads remain blocked. To address this limitation, we first develop a semantic-aware workload simulator that models deep learning jobs using intrinsic attributes such as model type, dataset size, and training configuration, and estimates execution behavior based on hardware-agnostic computational workloads. This enables more realistic modeling of heterogeneous performance and system-level behaviors. Building on this foundation, we propose SARL, a Semantic-Aware Reinforcement Learning scheduler that incorporates workload semantics and hardware heterogeneity into scheduling decisions. We formulate the scheduling problem as a Markov Decision Process and introduce a model-based GPU allocation mechanism to enable efficient decision making under large action spaces. Experimental results demonstrate that SARL consistently outperforms strong baselines, achieving a 29.6% reduction in deadline miss rate and a 14.4% reduction in job completion time (JCT).
  • Item type: Item , Access status: Open Access ,
    Towards Next-Generation Self-Sovereign Identity: Resilience, Privacy, and Trust in Decentralized Identity Systems
    (2026-07-24) Soltani, Reza; Nguyen, Uyen Trang
    Self-Sovereign Identity (SSI) has emerged as a promising paradigm for digital identity, shifting control of identity data from centralized intermediaries to individuals. However, its widespread adoption is impeded by open challenges in key management, fine-grained access control, regulatory alignment, and support for emerging Web 3.0 use cases. This dissertation investigates three overarching research questions: (1) How can users securely manage and recover their cryptographic keys in SSI applications in a way that is both privacy-preserving and usable? (2) How can users ensure that personal data is shared strictly according to their access preferences and consent? (3) Has SSI matured sufficiently to serve as a viable identity layer for contemporary digital ecosystems, including financial services and NFT-based platforms? To answer these questions, the dissertation combines conceptual modeling, systems design, and prototype implementation. First, it proposes and implements a practical key backup and recovery model for SSI wallets based on threshold secret sharing and decentralized key custodians, with optional biometric-derived factors. A prototype demonstrates that the scheme can improve resilience against key loss while limiting custodial trust and preserving user privacy. Second, the dissertation introduces a Data Capsule model that embeds user-defined access policies proposed by the users directly with encrypted data to control who and under what conditions the data can be access. The propose solution leverages decentralized identifiers, verifiable credentials, attribute-based encryption, and blockchain-based registries. This architecture enables selective disclosure, policy-based access, and auditable consent, and is evaluated in terms of performance and deployability on resource-constrained devices. Third, the dissertation designs SSI-based architectures for Know Your Customer (KYC) procedures and NFT authenticity in Web 3.0 environments. The KYC framework supports reusable, privacy-preserving credentials and reputation-based trust, while the NFT framework binds tokenized assets to verifiable identities and credentials, mitigating fraud and enhancing provenance. Through these use case specific designs the dissertation validates the maturity of SSI and its applications. Across these contributions, the dissertation offers a comprehensive survey of SSI foundations, platforms, privacy engineering techniques, and regulatory frameworks, and synthesizes their implications for real-world deployments. The results indicate that SSI can serve as a robust identity layer for diverse applications, provided that key recovery, access policies, governance, and user experience are carefully engineered.
  • Item type: Item , Access status: Open Access ,
    Toward Robust and Deployable Representation Learning for AI-Native Wireless Sensing
    (2026-07-24) Radwan, Ahmed Youssef; Tabassum, Hina
    Device-free Wi-Fi sensing offers a privacy-preserving approach to human activity recognition (HAR), yet practical deployment faces two key challenges: domain shifts across sensing environments and the gap between CSI compression and prediction in dynamic wireless systems. Existing Unsupervised Domain Adaptation (UDA) methods rely on labeled source data and fail in multi-user scenarios due to signal entanglement, permutation invariance, and severe class imbalance. AI-driven CSI feedback methods treat compression and prediction as independent problems, leaving channel aging insufficiently addressed. This thesis proposes two complementary frameworks. First, MU-SHOT-Fi is a Source-Free UDA framework for multi-user Wi-Fi sensing employing a permutation-invariant set-prediction architecture trained via Hungarian matching. It introduces an occupancy-weighted information maximization objective to handle class imbalance and integrates rotation-based spatial self-supervision to align cross-domain representations using the frequency-time structure of Channel State Information (CSI). A single-user extension, SU-SHOT-Fi, incorporates temporal self-supervision via Contrastive Predictive Coding (CPC). Second, a unified compression-prediction framework integrates CPC into a 3GPP-compliant CSI pipeline, jointly optimizing reconstruction fidelity and temporal predictive coherence to address channel aging without increasing feedback overhead. Two variants, CPC-before-Compression and CPC-after-Compression, offer different trade-offs between temporal modeling and device complexity. Evaluations on WiMANS and Widar 3.0 demonstrate consistent improvements over state-of-the-art baselines under cross-room and cross-frequency shifts. Experiments on Nokia, Oppo, and CATT datasets confirm favorable complexity-performance trade-offs, advancing AI-native wireless sensing toward robust, practical deployment.
  • Item type: Item , Access status: Open Access ,
    PARSE: A Framework for Evaluating Linguistic and Typographical Perturbation Bias in Large Language Models
    (2026-07-24) Lacalamita, John Pietro; Shvartzshnaider, Yan
    Large language models (LLMs) are increasingly becoming ubiquitous in day-to-day tasks. Yet, despite the growing dependence on LLM-based systems, their sensitivity to surface-level linguistic variation has received little attention. We introduce PARSE (Prompt Alteration Response-Shift Evaluation), a modular framework that generates grammatical, typographical, and dialectal prompt variants, queries LLMs under identical conditions, and measures distributional output shifts. We apply PARSE to two case studies, film recommendation and privacy bias evaluation, across three models (GPT-4o-mini, Llama~3.2, DeepSeek-7B). Results show that output shifts scale with perturbation intensity: grammatical rewrites produce minimal effects, while typographical noise and dialect rewrites significantly alter recommendations and appropriateness ratings. Perturbations push film recommendations toward higher-rated, generic titles, and shift privacy ratings toward more restrictive values. These effects are directionally consistent across models, demonstrating that the linguistic form of a prompt systematically biases LLM outputs.
  • Item type: Item , Access status: Open Access ,
    Connecting Text and Charts Using Large Vision-Language Models
    (2026-07-24) Chowdhury, Nafis Tahmid; Prince, Enamul Hoque
    Data visualizations are essential for presenting complex dataset, but the disconnect between charts and accompanying textual descriptions often leads to misinterpretation and increased cognitive effort—especially for users with limited data literacy. While prior methods attempt to bridge this gap, many depend on manual annotations or fixed chart structures, limiting scalability across diverse documents. In this thesis, we propose two large vision-language model (LVLM)-based frameworks—a single-agent baseline and a multi-agent architecture—for automatically linking textual descriptions with their corresponding chart data. Both frameworks extract structured data from chart images and use lexical, syntactic, and arithmetic reasoning to perform sentence-to-data alignment. We evaluate the performance of these frameworks on a curated dataset of Pew Research charts. Finally, we develop a browser extension that integrates this approach into Pew Research articles, enabling interactive text–chart linking for enhanced reading experiences.
  • Item type: Item , Access status: Open Access ,
    Inflammatory Biomarker Analysis from Wearable Sweat Patches via Smartphone-Based Image Processing
    (2026-03-10) Rozenblat, Shahak; Salahandish, Neda
    The detection of systemic inflammation through inflammatory biomarkers plays a critical role in identifying and managing pathological conditions. Conventional measurement of inflammatory biomarkers relies on invasive procedures such as blood sampling, which limits accessibility and requires frequent monitoring. Wearable sweat sensors offer a promising noninvasive alternative; however, robust interpretation of their visual signals remains challenging outside of laboratory environments. This study presents the first fully automated computational pipeline that translates colorimetric signals from a wearable sweat sensor into quantitative measurements of inflammatory biomarkers using smartphone-acquired images. The proposed approach enables reliable analysis under variable imaging conditions, supporting point-of-care (POC) inflammatory monitoring. Our results show that the pipeline significantly reduces measurement variability, achieving up to a 70\% reduction in variability that may be induced by different lighting conditions. Additional experiments demonstrate robustness across different smartphones and image capture distances with end-to-end processing completed within a few seconds. Furthermore, validation using data from human participants with eczema demonstrates that the system can distinguish between healthy individuals and those exhibiting elevated inflammatory biomarker levels, with performance comparable to the gold-standard of enzyme-linked immunosorbent assay (ELISA). The complete pipeline was integrated into a mobile application, enabling near real-time analysis and supporting practical POC deployment.
  • Item type: Item , Access status: Open Access ,
    PAMBA: Partition-aware and Multi-SLO Batching for Serverless Inference on Heterogeneous Clouds
    (2026-03-10) Abedini, Alireza; Khazaei, Hamzeh
    Serverless computing offers elasticity and fine-grained billing for machine learning inference, but efficiently supporting large models under diverse latency service-level objectives (SLOs) remains challenging. In particular, existing approaches face a wide cost–performance gap between CPU and GPU execution, while batching and resource selection become increasingly complex under heterogeneous workloads and multiple SLOs. This thesis presents PAMBA, a partition-aware and multi-SLO batching system for serverless inference on heterogeneous clouds. PAMBA combines multi-SLO batching with analytical latency and cost models for CPU, GPU, and partitioned execution, enabling consistent provisioning decisions across different execution modes. To bridge the CPU–GPU gap, the system employs a customized partitioning strategy derived from latency-optimal partitioning, adapted to satisfy serverless resource constraints and jointly consider latency feasibility and per-request cost. This adaptation allows partitioned execution to emerge as an effective intermediate regime between monolithic CPU and GPU deployments. By jointly optimizing execution mode selection, batching, and resource allocation, PAMBA enables flexible inference deployment across a wide range of SLOs and arrival rates, including scenarios where GPU resources are unavailable or inefficiently utilized. Experimental results on convolutional neural networks demonstrate that PAMBA identifies distinct execution frontiers and reduces inference cost compared to existing serverless batching techniques, while maintaining SLO feasibility across heterogeneous workloads.
  • Item type: Item , Access status: Open Access ,
    OptiServe: Cost-Aware, Performance-Driven, and Accuracy-Tuned Serverless Applications with ML Workloads
    (2026-03-10) Boukani, Arian; Khazaei, Hamzeh
    Serverless computing has emerged as a popular cloud paradigm due to its seamless scalability and cost-efficient, pay-as-you-go pricing model. Its potential to support machine learning (ML) inference workloads,including generative AI tasks, has led to growing adoption of ML functions within serverless applications. A key challenge, however, is selecting suitable ML models that balance execution time, deployment cost, and inference accuracy in latency- and cost-sensitive environments. In this study, we present a framework for optimizing serverless applications that incorporate ML components through tri-objective optimization. We develop high-fidelity analytical models, augmented with lightweight profiling, to capture the trade-offs among cost, performance, and accuracy across different model choices. These models serve as the foundation for guiding ML model selection and deployment strategies to meet application-specific service-level objectives. We validate our framework through real-world experiments on AWS using real serverless applications. Furthermore, we demonstrate its practicality by performing extensive what-if analyses, exploring a wide range of application scenarios and configurations, in under a minute. Our extensive experiments on real-world applications show that OptiServe recommends memory and ML model configurations that achieve over 95% of the accuracy of ideal configurations in 89.64% of cases, enabling efficient, low-cost deployments while maintaining model accuracy and meeting performance targets.
  • Item type: Item , Access status: Open Access ,
    Foundation Models for Analyzing Single-Cell RNA Sequence data
    (2026-03-10) Naziri, Amirreza; Seyyed-Kalantari, Laleh
    Single-cell RNA sequencing (scRNA-seq) measures gene expression in individual cells, offering deep insight into cellular heterogeneity, development, and disease. Transformer-based foundation models have become central to single-cell RNA-sequencing analysis, yet most rely on uniform random masking during pretraining, a strategy misaligned with the sparsity, heterogeneity, and zero inflation characteristic of scRNA-seq data. To assess how these models behave under realistic biological variation, we first perform a comprehensive evaluation of four widely used single-cell foundation models (Geneformer, scBERT, scFoundation, and scGPT) across three diverse datasets. This benchmarking reveals substantial variability in model performance, including systematic weaknesses on rare cell populations and degraded accuracy in clinically challenging conditions. Motivated by the broader limitations of random masking in Foundation models, we introduce Multinomial Attention Masking (MAM), a biologically informed masking strategy that leverages trainable latent representations and cross-attention to identify informative gene positions during pretraining. Across all datasets, models pretrained with MAM consistently achieve higher downstream cell-type classification accuracy than those trained with uniform masking and, in several cases, outperform the original pretrained backbones. Biological validation further demonstrates that MAM preferentially selects highly expressed and functionally meaningful genes, indicating that its improvements stem from capturing biologically relevant structure rather than from increased algorithmic complexity. This work improves the reliability and utility of single-cell foundation models for researchers and clinicians alike.
  • Item type: Item , Access status: Open Access ,
    Towards Agentic Vision Language Models for Question Answering on Interactive Dashboard
    (2026-03-10) Kartha, Aaryaman Sudhir; Prince, Enamul Hoque
    Multimodal models, specifically Vision Language Models (VLMs), have shown increasing capabilities in data visualization oriented downstream tasks, achieving performance saturation in shorter intervals of time. Consequently, focus has shifted to assessing their potential towards new frontiers, specifically interactive environments. Various benchmarks center around data visualization question answering tasks on static visualizations, and such rudimentary approaches don’t reflect real world analysis scenarios where vast decision making is required. Dashboards, while being commonplace tools in various industries, have had limited work done into evaluating the capabilities of VLMs to traverse and reason with them. To tackle these limitations, this thesis presents DashboardQA, a novel benchmark for interactive dashboard question answering. Overall, 292 tasks encompassing 405 QA pairs are presented from 5 diverse category types, with 112 carefully chosen dashboards represented. Experimental results show this benchmark is a challenge for various types of VLMs assessed, with the best model achieving 38.69 %.
  • Item type: Item , Access status: Open Access ,
    Designing an Interactive Tool for Mnemonics Creation and Knowledge Retention
    (2026-03-10) Ejaz, Sarah; Oyibo, Kiemute
    Research shows that mnemonics are an effective learning technique, yet few tools support mnemonics-based long-term learning. We designed and evaluated a mnemonics-creation tool, the SAVE Tool, to promote active learning and retrieval practice. Forty-five participants were assigned to experimental and control groups and viewed a 10-minute biology lecture covering six topics. They completed recall and recognition tasks after a 45-minute practice session (T1), one week later without revision (T2), and after a 15-minute revision (T3). Results showed that the SAVE Tool group consistently outperformed the control group in recall across all time points, with statistically significant differences at T3 and for more difficult topics such as Krebs Cycle Substrates and Cranial Nerves. No significant group differences were found for recognition. These findings suggest mnemonics-based tools can enhance long-term learning without hindering understanding and should be integrated into memory-intensive courses.
  • Item type: Item , Access status: Open Access ,
    Designing Mnemonics Serious Games to Promote Knowledge Retention in Memory-Intensive Courses
    (2026-03-10) Fung, Kingson; Oyibo, Kiemute
    The increased difficulty of memory-intensive courses due to many factors necessitates using technological tools to promote retrieval practice and long-term learning. Hence, a mnemonics game, based on the RADAR framework proposed by Oyibo, was implemented to foster knowledge retention. Fifty-two students, comprising an experimental group (n = 30) and a control group (n = 22), were recruited to undertake a study, which involved watching a 10-minute Biology lecture on Biology Organization, Cranial Nerves, and Krebs Cycle and taking repeated tests immediately after a 45-minute preparation, one-week gap pre- and post-15-minute revision. In all three tests, students who used the RADAR game performed better in recall than those who did not. Coupled with the study participants stating they found the game easy to use, enjoyable, useful, and trustworthy, and their willingness to adopt it, the experimental group’s better performance highlights the need to incorporate mnemonics-based games in memory-intensive courses to promote long-term learning.
  • Item type: Item , Access status: Open Access ,
    A Concurrent List with Adaptive Bounds
    (2026-03-10) Asbell, Shalom Moshe; Ruppert, Eric
    Few concurrent data structures adapt dynamically to access patterns, and those that do lack formal performance guarantees. This thesis introduces the first self-adjusting concurrent data structure with an adaptive bound analogous to those of optimal self-adjusting sequential structures. We present a lock-free move-to-front (MTF) list supporting a dynamic set of keys, where an operation op on a key k runs in an amortized number of steps proportional to the size of its working plus a contention term. The working set of op is the set of keys accessed since the last operation on key k, and contention is the number of operations that run concurrently with op. We further show that the list performs at most twice as much work as the best possible fixed list for a set of searches, up to contention. Finally, we prove that the number of nodes reachable from shared memory is bounded by the number of keys in the set plus contention.
  • Item type: Item , Access status: Open Access ,
    "Towards Closed-Loop Sleep Monitoring in Parkinson’s Disease: Self-Supervised Learning Strategies for Sleep Stage Classification"
    (2026-03-10) Menguc, Kristal Doga; Zylberberg, Joel
    Parkinson’s disease (PD) involves severe sleep disturbances that may accelerate neurodegeneration. Closed-loop deep brain stimulation (DBS) is a promising therapeutic solution but requires accurate, real-time sleep-stage classification from subthalamic nucleus signals, a task where conventional models generalize poorly. To address pronounced class imbalance and improve cross-patient generalization, this work introduces a self-supervised transformer framework. The model employs a masked autoencoder strategy, pretrained on large public EEG/ECoG datasets from healthy subjects to learn balanced representations of all sleep stages, thereby improving discrimination of rare classes. While pretraining yielded limited overall improvement, it specifically enhanced N3 stage identification and next-sleep-stage prediction. Contrastive self-supervision (70%) significantly outperformed a reconstruction-based approach (62%). Furthermore, spectral feature extraction proved more effective than temporal CNN features for distinguishing commonly mispredicted stages. Future work will focus on hybrid reconstruction-contrastive losses, incorporating spectral feature usage to forecasting and extending the forecasting horizon to predict multiple subsequent stages for proactive neuromodulation.
  • Item type: Item , Access status: Open Access ,
    A CNN–LSTM–Attention Hybrid Architecture for Real-Time Intrusion Detection at the Data Link Layer
    (2026-03-10) Ahmadnejad Roudsari, Amirhossein; Habibi Lashkari, Arash
    Data Link Layer (Layer 2) security remains one of the most underexplored areas in modern network intrusion detection research, despite its critical role as the foundation of reliable communication between networked devices. Attacks at this layer, such as ARP spoofing, MAC flooding, VLAN hopping, and DHCP starvation, can compromise entire networks before higher-layer defenses activate. Existing intrusion detection systems predominantly focus on network or transport layers, leaving a significant gap in early-stage threat prevention. To address this limitation, this thesis proposes a memory-efficient hybrid deep learning architecture that integrates Convolutional Neural Networks (CNNs), Long Short-Term Memory (LSTM) units, and an Attention mechanism for real-time detection of Layer 2 intrusions. A novel dataset, BCCC-DLLayer-IDS-2025, was developed as part of this research, comprising over 4.6 million labeled flow records collected in a controlled experimental environment. The dataset includes eleven distinct attack types spanning spoofing, flooding, and protocol manipulation scenarios, along with benign traffic, providing a comprehensive foundation for training and benchmarking Layer 2 intrusion detection systems. The proposed CNN–LSTM–Attention architecture combines spatial and temporal feature extraction with an adaptive focus mechanism, enabling effective modeling of short-term dependencies in network traffic while reducing redundancy. The model achieves an F1-score of 99.67\% with only 2.1 million parameters and a latency below 100 milliseconds, offering a 60\% lower computational cost than conventional deep learning models. Extensive experiments under varying traffic conditions and noise levels confirm the model’s robustness, generalizability, and suitability for real-time deployment on resource-constrained edge and IoT devices.
  • Item type: Item , Access status: Open Access ,
    Monocular Camera-based Road Segmentation Using Geometric Cues
    (2026-03-10) Cheng, Gong; Elder, James
    Vision-based road segmentation aims to identify drivable road regions in images captured by vehicle-mounted cameras. With the rapid development of autonomous driving, such methods have attracted substantial attention from both academia and industry due to their research significance and commercial potential. However, a key challenge lies in achieving robust adaptation—that is, effectively generalizing to diverse environmental and operational conditions. In my Ph.D. research, I have concentrated on leveraging geometric information to advance monocular-camera-based road segmentation and adaptation. Specifically, I have completed three independent projects: (1) a supervised road segmentation approach that fuses geometric cues with appearance cues to enhance segmentation accuracy, (2) an unsupervised domain adaptation method that employs a novel geometry-guided strategy to reduce domain shift between source and target data, and (3) another unsupervised domain adaptation approach that uses known camera parameters and online estimation of the camera pan and tilt to better align imagery across and within datasets, improving accuracy and generalization. These contributions collectively push forward the development of robust, geometry-driven solutions for road scene understanding.