Information Systems and Technology

Permanent URI for this collectionhttps://hdl.handle.net/10315/27588

Browse

Recent Submissions

Now showing 1 - 20 of 83
  • Item type: Item , Access status: Open Access ,
    A Novel Computational Ethology Framework for Studying Animal Behaviour Under Climate Change (CEFABC2): A Case Study of Little Penguins on Phillip Island
    (2026-07-24) Camelo Guerrero, Abraham Isaac; Khaiter, Peter A.
    Animal behaviour is a biological indicator that provides insight into a species’ interaction with its environment. Long-term records of animal behaviour form time series containing rhythms, anomalies, and trends in the activity of a species. Numerical and machine learning (ML) methods provide tools to automate the investigation of large amounts of data to detect patterns in complex behavioural time series. This thesis uses the little penguin colony on Phillip Island, Victoria, Australia as a case study, where the behaviour of this species has been recorded and analyzed over decades. To further support these studies, this thesis introduces a novel computational ethology framework on the study of animal behaviour under climate change (CEFABC2). This framework extracts and forecasts the dynamics of nightly penguin count time series, breeding success, and mean egg-laying date (MLD). CEFABC2 is built on four modules: the first module uses singular spectrum analysis (SSA) to process and reconstruct the nightly penguin count time series. From the SSA-reconstructed time series, the peaks are extracted and Gaussian Mixture Model (GMM) is implemented to estimate the peak width. The second module examines the correlations between the extracted peaks, the MLD, and breeding success. Gaussian process (GP) is applied to forecast the peaks of the nightly penguin count time series. Additionally, the performance of structural time series (STS), ARIMA, LSTM, GP, LightGBM, and Ridge models is evaluated for forecasting nightly penguin counts. The third module describes the climate of Phillip Island, using principal component analysis (PCA), K-means, and GMM to identify meteorological seasons. The fourth module integrates biological and climate variables, analyzing their correlations and forecasting breeding variables using ElasticNet, Ridge, Random Forest, and Gradient Boosting ML models. Finally, the CEFABC2 management system incorporates these components into a single tool aiming to support monitoring, forecasting, and future conservation management efforts towards the Phillip Island little penguin colony.
  • Item type: Item , Access status: Open Access ,
    Interpretable Deep Tabular Learning for Fraud and Phishing Detection in Decentralized Finance (DeFi)
    (2026-07-24) Ameri, Ava; Habibi Lashkari, Arash
    Decentralized Finance (DeFi) has introduced new security challenges due to its open, permissionless, and pseudonymous nature, which has increased the risk of fraud and phishing activities. This thesis first presents a comprehensive study of 284 DeFi platforms to examine their architectural, functional, and security-related characteristics. Building on this ecosystem-level analysis, the thesis develops a behavior-centric multiclass detection framework using Ethereum transaction data. The framework integrates legitimate, fraud, and phishing activities into a unified dataset and evaluates several traditional and deep tabular learning models, including TabNet, GANDALF, and NODE. The results show that deep tabular models outperform conventional baselines, with NODE achieving the strongest overall performance. Feature importance analysis highlights gas usage, nonce behavior, transaction frequency, and wallet activity as key indicators of malicious behavior. Overall, this thesis demonstrates that behavior-based Ethereum transaction features combined with deep tabular learning can support more effective and scalable DeFi threat detection.
  • Item type: Item , Access status: Open Access ,
    Understanding Worker Perception of Pay Fairness on Microtask Crowdsourcing Platforms
    (2026-07-24) Liu, Yiduo; Jiang, Ling
    This thesis investigates the factors shaping worker perception of pay fairness on microtask crowdsourcing platforms. It employs a hybrid approach that integrates data-driven theory building with theory-driven sensemaking and empirical testing informed by organizational justice theory. Collecting and analyzing 14,553 reviews posted by Amazon Mechanical Turk workers on TurkerView, the study identifies seven latent themes each for positive and negative experiences and maps them onto distributive, procedural, and interactional justice dimensions. Distributive concerns (e.g., generous pay, bonus opportunities, easy earnings, and time efficiency) emerge as salient factors, while procedural and interactional factors reflect task design quality and requester communication. Predictive models, including stepwise econometric models and a DistilRoBERTa-based deep learning model, show that dense text embeddings capture fairness-relevant semantic information beyond topic-level abstractions in explaining pay fairness ratings. The findings advance understanding of pay fairness and inform platform design, requester practices, and labor policy in the gig economy.
  • Item type: Item , Access status: Open Access ,
    ASC-PIE: An Evaluation Framework for PII-Aware Named-Entity Recognition
    (2026-07-24) Hafez, Mohamed Mohamed Nader; Litoiu, Marin
    The robust extraction of Personally Identifiable Information (PII) is essential for privacy protection in modern text-processing systems, where users often share sensitive details in prompts, emails, chat logs, and support tickets. As PII categories and deployment domains evolve, updating extraction models can improve coverage but may also cause catastrophic forgetting of previously learned types. This thesis investigates how PII extraction can remain accurate, reliable, and maintainable as task scope expands. It introduces ASC-PIE, a unified English corpus and evaluation framework that combines public datasets with a synthetic component to improve coverage of rare and challenging PII cases without using real personal data. Using ASC-PIE, the thesis compares encoder-based, encoder-decoder, and decoder-only models under supervised fine-tuning and in-context prompting. It also proposes SPRINT-PP, a privacy-safe continual-learning method that mitigates forgetting without storing raw historical examples. The evaluation covers extraction quality, robustness, output validity, computational efficiency, and knowledge retention.
  • Item type: Item , Access status: Open Access ,
    Deep Learning Framework for High-Resolution Hydrological Forecasting Using Hydrometric Data
    (2026-07-24) Rahman, Md Asifur; Erechtchoukova, Marina G.
    Multi-site multi-horizon hydrological forecasting is gaining increasing attention in watershed management due to spatially distributed hydrological processes, which cannot be represented using single site models. This study develops a deep learning (DL) framework for forecasting hydrological regime in streams at multiple cross-sections using hydrometric observations. It has been shown that multi-site models with a shared static loss function better represent the sites with regular hydrological regimes and moderate flows, while rapid response sites are not modelled accurately. To address this limitation, an Adaptive Weighted Loss strategy was developed to periodically re-assign weights to site losses depending on the validation performance. The framework was assessed on datasets with different hydrological characteristics. The results highlighted that ALF-based DL framework improved performance at response sites by extending horizon of reliable forecasts. Overall, the study shows that combining watershed-level multi-site forecasting with adaptive weighting provides a scalable and reproducible approach for high-resolution hydrological predictions.
  • Item type: Item , Access status: Open Access ,
    Optimizing Urban Safety and Traffic Management: A Machine Learning Approach to Sensor Placement for Intelligent Transportation and Crime Detection Systems
    (2026-07-24) Denis Nedeljkovic; Jammal, Manar
    This thesis explores the use of machine learning algorithms for optimizing sensor placement in urban areas to improve crime prevention and traffic management. Focused on the City of Toronto as a case study, it integrates traffic sensor data with vehicular crime statistics to propose a model predicting potential hotspots for traffic violations and optimal locations for crime-prevention sensors. This research employs the use of Random Forest, Long Short-Term Memory, Fourier Series Neural Networks, Support Vector Machine, and the Feed Forward Neural Network, to provide insights that are actionable for safer, smart cities. Providing these results for government officials, law enforcement agencies, and research with an ease of access cloud-based tool to serve as a Software as a Service format. The study underscores the importance of leveraging advanced data analytics in urban planning, suggesting a direction for future research and implementation in intelligent transportation systems.
  • Item type: Item , Access status: Open Access ,
    "Ethics, Governance, and Bias Mitigation in AI-Enabled loT Applications: A Focus on Transparency and Fairness
    (2026-07-24) Nadine Fares; Jammal, Manar
    In the development of smart cities, the integration of Internet of Things (IoT) and Artificial Intelligence (AI) technologies promises to address urban challenges yet raises significant security and ethical concerns. This study emphasizes the need for robust governance frameworks to ensure the ethical deployment of these technologies, focusing on the roles of data privacy, security, and bias mitigation. It highlights the importance of Explainable AI (XAI) techniques, like LIME and SHAP, for enhancing transparency and accountability in AI decision-making processes. Through a detailed examination of existing governance frameworks and the introduction of a bias detection tool, this work aims to embed ethical considerations into the fabric of AI-driven systems in smart cities. The discussion extends to exploring common domains and AI implementations across different industries, offering insights into comprehending and ensuring the responsible use of AI and IoT technology. This lays the groundwork for future research in this developing area.
  • Item type: Item , Access status: Open Access ,
    AI-Driven Fake News Detection: Trends, Techniques, and Experimental Analysis
    (2026-03-10) Zeng, Li; Huang, Jimmy
    The rapid spread of fake news poses a significant challenge to information accuracy. This thesis highlights fake news definitions and characteristics, introducing a taxonomy that categorizes AI-driven detection methods into model-centric and process-centric approaches. We evaluate various approaches ranging from traditional machine learning to trending AI methodologies, focusing on techniques like data augmentation, information extraction, and results explanation. To re-evaluate classical algorithms, this work provides a detailed analysis of a 2016 U.S. election dataset. By employing fact-checking and advanced data mining, we investigate linguistic characteristics through exploratory data analysis and apply multiple machine learning algorithms for classification. Experimental results yield valuable insights into the defining characteristics of fake news and demonstrate machine learning's potential to enhance misinformation filtering. Finally, we discuss four main challenges and trends aimed at refining detection accuracy and integrating cutting-edge AI methodologies to combat fake news more effectively.
  • Item type: Item , Access status: Open Access ,
    Efficient Text-Image Retrieval Using Large Language Models
    (2026-03-10) Liu, Jiahao; Yu, Xiaohui
    Efficient retrieval from large-scale image databases is a key challenge, particularly as applications increasingly rely on multimodal models such as CLIP. While CLIP offers strong joint image–text representations for semantic search, its globally pooled embeddings often struggle with fine-grained, multi-concept queries, leading to high false positives and reliance on costly verification models. To address this, we propose a hybrid framework that structures the embedding space through feature clustering and models candidate selection as a multi-armed bandit problem. Each cluster acts as an arm, with relevance scores from ground-truth systems as rewards. Using Thompson Sampling, this approach balances exploration and exploitation to quickly identify promising clusters, reducing unnecessary ground-truth queries. Experiments show that our method significantly improves precision and lowers computational costs in multi-keyword retrieval tasks, enabling scalable, fine-grained retrieval in resource-constrained settings. This structured, adaptive approach effectively enhances CLIP-based retrieval pipelines.
  • Item type: Item , Access status: Open Access ,
    Query Generation for Database Testing Via Machine Learning
    (2026-03-10) Yang, Yongtai; Yu, Xiaohui
    Modern database management systems (DBMSs) are important to data-driven applications. However, testing DBMS bugs is still a challenging task as the DBMS is a very complex system. Bugs in the DBMS often appear only under specific execution plan patterns, such as nested-loop joins combined with aggregation. Reproducing such bugs requires generating SQL queries whose execution plans contain the pattern that triggers the bug. Existing rule-based query generators and learning-based approaches both fail to generate queries under the execution plan pattern constraint. To overcome this limitation, we propose QueryMorpher, a plan-driven query generation framework that generates SQL queries from the problematic execution plan that triggers the bug. QueryMorpher begins with a problematic execution plan and a plan pattern that triggers the bug, and implements a sequence of learned plan mutation operations guided by a sequence-to-sequence model. The mutated plan is then translated back into SQL by using a plan-to-query translation module, which guarantees that the resulting query reproduces the desired execution plan while remaining syntactically and semantically valid. Experimental results demonstrate that QueryMorpher can generate diverse and valid queries whose execution plans contain the user-defined patterns. On TPC-H, QueryMorpher achieves a target-pattern rate of 0.6 vs 0.4 for the best baseline, while maintaining 10% higher plan diversity under the same budget. On TPC-DS, QueryMorpher achieves similar improvements, indicating that QueryMorpher is stable on different database schemas. By bridging the gap between query generation and query execution plan control, QueryMorpher enables automated and controllable DBMS testing.
  • Item type: Item , Access status: Open Access ,
    Counterfactual Prescriptions Via Hierarchical ML For Missed Chemotherapy Appointment Prevention
    (2026-03-10) Rajabzadeh, Mohammadreza; Senderovich, Arik
    Missed medical appointments, including cancellations and no-shows, disrupt clinical workflows, reduce efficiency, and compromise patient care. Using 1.8 million chemotherapy appointments from the Dana-Farber Cancer Institute, this study develops a hierarchical machine learning framework that first predicts cancellations and then no-shows among remaining cases. Combining operational and temporal features, the model achieves F1-scores of 0.76 and 0.82, outperforming a multinomial baseline by 7–10 points on minority classes. To extend interpretability, we infer missed-visit reasons through semi-supervised learning on short-notice cancellations, training supervised predictors with weighted F1-scores of 0.57 and 0.54. Counterfactual analysis identifies modifiable scheduling factors—such as appointment timing and provider consistency—while showing that standard interventions like reminders are less effective than previously reported. These findings underscore the value of predictive models with prescriptive intent, highlighting their role in designing tailored, context-specific interventions to improve clinical efficiency and patient care.
  • Item type: Item , Access status: Open Access ,
    Integrating Natural Language Processing with Expert Systems for Streamlined Evaluation of Applications
    (2025-11-11) Mohanty, Arup; Khaiter, Peter A.
    Screening of applications for jobs, education, grants, and the like can often be slow, subjective, and inefficient. To address these issues, a novel framework, the Hybrid AI Platform for Streamlining Evaluation (HAIPSE), is introduced, integrating expert rule-based reasoning with NLP and computer-vision techniques, delivering a structured alternative to traditional applicant tracking systems. The platform’s heuristic scoring module, powered by spaCy, extracts key application details and grades responses against predefined criteria. To capture context and reduce reviewers’ workload, the HAIPSE incorporates large language models, Meta Llama 3-8B and Mistral 8 × 7B, generating concise essay summaries. The novel group-fairness metrics are applied within the evaluation pipeline, making the scoring process more transparent while preserving nuanced content. Furthermore, built-in debiasing steps embed proactive fairness checks directly into the framework’s design, preventing bias rather than merely detecting it post factum. All components of the HAIPSE were trained and tested on a limited set of sample IDs and 2,000+ real applications from the NIB Trust Fund (Canada). Compared with manual reviews, the HAIPSE improves transparency and reduces bias, while a collaborative audit interface bridges AI automation and expert judgment, reinforcing responsible AI. Ethical considerations of fairness and responsible deployment ensure that the resulting assessments remain scalable, interpretable, and equitable.
  • Item type: Item , Access status: Open Access ,
    A Systematic Evaluation Framework for Smart Contract Security Analyzers: Methods, Metrics, and Framework
    (2025-11-11) Hejazi, Niosha; Arash Habibi Lashkari
    Smart contracts automate agreements in blockchain systems but their immutable nature makes them vulnerable to permanent flaws once deployed. This thesis evaluates 256 smart contract vulnerability detection tools developed between 2018 and 2024, including approaches such as fuzzing, symbolic execution, formal verification, and artificial intelligence–based analysis. Tools were classified by detection strategy (static, dynamic, hybrid), domain (academic or industry), and scope. The evaluation involved a theoretical review of architecture, usability, and documentation, alongside an empirical assessment of accuracy, speed, and false positive rates. Findings show that while certain tools excel in specific areas, none achieve balanced performance or comprehensive coverage. To address these gaps, a modular six-layer evaluation framework is introduced, defining functional areas such as code analysis, coverage, integration, and user experience. The framework offers a benchmark for tool assessment and future development. Additionally, a graph-based detection model is proposed, demonstrating improved accuracy in both binary and multi-class settings.
  • Item type: Item , Access status: Open Access ,
    An MQTT Application-Layer Traffic Analyzer for Interpretable Flow-Level Intrusion Detection and Zero-DayThreat Identification in IoT Environment Using TabNet
    (2025-11-11) KouhiRonaghi, Arefeh; Habibi Lashkari, Arash
    Message Queuing Telemetry Transport (MQTT) is widely used in IoT systems; however, its lightweight design makes it vulnerable to various cyberattacks. This research reviews existing intrusion detection methods for MQTT and shows their limitations in detecting new and complex threats. This study presents a comprehensive intrusion detection framework that uses raw PCAP data and employs flow-based behavioral analysis to detect both known and novel attacks. We present MQTTFlowLyzer, a protocol-aware analyzer designed to extract detailed MQTT flow features and generate an augmented dataset, BCCC-MQTT-IDS-2025, that captures realistic and diverse attack scenarios. The extracted features train a TabNet-based learning model capable of integrated feature selection, classification, and confidence-based detection of zero-day threats. Our approach highlights the behavioral uniqueness of each attack class and uses attention-driven interpretability for in-depth analysis. Experimental results demonstrate that the model effectively detects attacks while maintaining high performance across other categories. The system successfully flags previously unseen traffic by profiling class-specific behaviors and incorporating confidence thresholds. These results demonstrate the potential of flow-based, interpretable learning for real-time and resilient MQTT intrusion detection.
  • Item type: Item , Access status: Open Access ,
    Selective Cloud Offloading for Accurate and Efficient Object Detection
    (2025-11-11) Dehghani Firoozabadi, Davood; Yu, Xiaohui
    High-accuracy object detection on resource-constrained devices is essential for many applications including autonomous systems, agriculture, and mobile computing. However, deploying high-performance object detection models on these devices is impractical due to computational limitations, and transmitting and processing all data on a much more powerful remote server running significantly more complex and accurate models, known as full cloud offloading, incurs high latency and cost. This thesis proposes a selective cloud offloading framework that balances prediction accuracy and processing cost. A lightweight edge model makes initial predictions using conformal prediction to quantify uncertainty. Only high-uncertainty regions are offloaded to the cloud for refinement by more powerful models. To further optimize efficiency, multiple uncertain regions are merged into a single image before offloading, reducing transmission and processing costs. The system is evaluated on real datasets, demonstrating substantial accuracy improvements with minimal additional overhead.
  • Item type: Item , Access status: Open Access ,
    Developing Advanced Representation Learning Techniques for mRNA Sequence and Structure Modeling
    (2025-11-11) Nahali, Sepideh; Huang, Jimmy
    Recent studies in bioinformatics and genomics focus on analyzing RNA sequences, which are complex due to diverse nucleotide compositions, varying lengths, and multiple isoforms. Accurately modeling these sequences is essential for predicting mRNA degradation, a key factor in designing effective RNA-based therapies. However, many existing models struggle to capture the intricate relationships between sequence and structure, limiting their predictive power. We introduce StructmRNA, a BERT-based model using dual-level and conditional masking to embed RNA sequences and structures. This enables accurate prediction of mRNA sequences and structures without explicit structural data, effectively capturing sequence-structure dependencies. Evaluations show StructmRNA outperforms existing models in predicting mRNA degradation and secondary structure. Experiments with GAN-generated RNA sequences showed no performance improvement. Nonetheless, StructmRNA’s consistent convergence over 30 epochs highlights its robustness and accuracy. This work advances RNA representation learning and demonstrates deep learning’s potential in RNA-based therapeutic design and bioinformatics.
  • Item type: Item , Access status: Open Access ,
    Characterizing Osteosarcopenia In Spinal Metastases Patients Undergoing Stereotactic Body Radiotherapy (Sbrt): Leveraging Deep Learning For Improved Outcome Prediction
    (2025-07-23) Castano Sainz, Yessica Caridad; Chen, Stephen
    Stereotactic body radiotherapy (SBRT) is commonly used to treat spinal metastases, offering excellent local control and pain relief. However, it carries an average 14% risk of vertebral compression fractures (VCFs), and despite growing evidence linking osteosarcopenia to adverse clinical outcomes, musculoskeletal health is not routinely assessed during SBRT treatment planning. This thesis introduces a fully automated pipeline for extracting musculoskeletal biomarkers from CT, combining deep learning–based segmentation with vertebral landmark–guided cropping and volumetric analysis. Sarcopenia thresholds were derived for volumetric indices using height-based and vertebral-based normalization, guided by established literature cutoffs for the Psoas Muscle Index (PMI). Osteoporosis was defined using trabecular bone density. In this SBRT cohort, 58% of patients met criteria for osteoporosis, 45% for sarcopenia, and 31% for osteosarcopenia. In multivariable logistic regression analyses, significant associations between fracture risk and both osteoporosis and lower psoas muscle density were observed in specific models, warranting further investigation. Additionally, categorical definitions of sarcopenia and osteosarcopenia were significantly associated with reduced overall survival. The pipeline was extended to MRI using CT-based segmentations as weak labels for training nnU-Net models, achieving high segmentation accuracy and supporting future radiation-free musculoskeletal biomarker assessment.
  • Item type: Item , Access status: Open Access ,
    Grasp - A Graph-Based SLA Breach Prediction Framework at the Service Level in Neural Inference
    (2025-07-23) Fehresti, Sara; Litoiu, Marin
    Cloud computing, as the backbone of modern adaptive software architectures, has revolutionized data storage and processing, driven by the power and flexibility of microservices. Despite their advantages in fault isolation and flexible deployment, microservices often experience unpredictable latency spikes, leading to costly Service Level Agreement (SLA) violations. This thesis introduces a multiscale time-spectrum framework, called GRASP (Graph-based SLA Breach Prediction). It leverages time-series data, sequence processing, and graph-based modeling to proactively detect performance anomalies and predict SLA breaches in microservice-based systems within upcoming time windows. In our framework, raw data is transformed into graph representations and fed into deep learning models to capture both topological and temporal characteristics. By combining graph analysis with sequential modeling, our dual approach not only identifies critical service dependencies but also pinpoints potential end-to-end bottlenecks. Evaluations on microservice datasets demonstrate its superiority over baseline methods in early warnings, forecasting breaches, and localizing root causes at the service level, underscoring its potential to enhance the reliability and efficiency of microservice-based applications in cloud environments.
  • Item type: Item , Access status: Open Access ,
    Service-Level Prediction And Anomaly Detection Towards System Failure Root Cause Explainability In Microservice Applications
    (2025-07-23) Rouf, Raphael; Litoiu, Marin
    Microservice and cloud computing operations are increasingly adopting automation. The importance of models in fostering resilient and efficient adaptive architectures is central to ensuring that services operate with expected behavior and performance. To effectively predict, detect, and explain system failures, it is essential to develop a comprehensive understanding of the affected application. This means examining its unexpected behaviors from multiple perspectives, including logs, metrics, and dependencies, to uncover the underlying root causes. This thesis presents a novel approach to system failure prediction, root cause analysis and explainable failure type analysis by leveraging a three-fold modality of IT observability data: logs, metrics, and traces. The proposed methodology integrates Graph Neural Networks (GNN) to capture spatial information and Gated Recurrent Units (GRU) to encapsulate the temporal aspects within the data. A key emphasis lies in utilizing a stitched representation derived from logs, microservice events and resource metrics to predict system failures proactively. The traces are aggregated to construct a comprehensive service call flow graph and represented as a dynamic graph. Furthermore, permutation testing is applied to harness node scores, aiding in the identification of root causes behind these failures. We evaluate our approach on open source datasets: MicroSS, QOTD and Train Ticket dataset that captures various types of system faults such as resource overload and wrong manipulation faults. Our findings on real world cases demonstrates that the proposed three-fold modality of observability data based on enhanced preprocessing and applying GNN-GRU and gradient distance explanation model captures failure type predictions well and explains them effectively for engineers to help debug and diagnose the issue.
  • Item type: Item , Access status: Open Access ,
    Application of Remote Sensing and Machine Learning in Vegetation Phenology and Climate Change Studies
    (2025-07-23) Suleman, Masooma Ali Raza; Khaiter, Peter A.
    Remote sensing and machine learning (ML) have revolutionized phenology studies by offering scalable and automated methods for monitoring vegetation growth patterns. Traditional phenology detection methods, which rely on field observations, are often labor-intensive and geographically constrained. This thesis introduces a novel deep learning model, the Temporal Multivariate Attention Network (TMANet), which integrates remote sensing data, climate indices, and ground observations to enhance phenological stage detection in crops. Focusing on corn phenology, the study explores how remote sensing data preprocessing optimizes its utility for phenology applications, how ML techniques improve detection accuracy, and how TMANet outperforms traditional models in capturing temporal and environmental dependencies. The proposed framework provides a robust, data-driven approach to understanding vegetation responses to climate variability, supporting sustainable agricultural management. The findings contribute to advancing phenology research by offering a scalable and efficient methodology for monitoring crop development and assessing climate change impacts on vegetation phenology.