About

I am a Data Science student at the University of Melbourne with an interest in risk, governance, and assurance. I am particularly interested in problems where complex systems need to be understood carefully, whether through modelling, analysis, or control evaluation.

My background spans statistical modelling, financial and investment risk, technology risk, governance, and assurance. What connects these areas is an interest in understanding how reliable a system or decision remains when conditions change, relationships between variables shift, or different parts of a system begin to behave differently.

Across my work, a consistent focus is understanding how models, signals, controls, and systems behave in practice, and what this means for the decisions that depend on them.

Projects

Credit Risk Under Regime Change

Python | Statistical Modelling | Time Series Evaluation

This project investigates how credit risk models behave when the environment they were trained on changes. Using LendingClub loan data from 2007 to 2015, the objective was not only to predict default, but to examine how model reliability changes when relationships between borrower characteristics and outcomes shift over time.

Logistic regression models were evaluated with and without regularisation, while random forest hyperparameters were tuned using Bayesian optimisation. Both approaches used borrower and loan characteristics including credit grade, interest rate, debt-to-income ratio, income, loan amount, and term. Rather than relying on a conventional random train-test split, performance was evaluated through a rolling time-based framework, with models trained on recent historical data and tested over progressively later periods.

Key components include:

  • Designing a time-aware evaluation framework to assess model reliability under distribution shift
  • Comparing linear and tree-based models in terms of predictive performance, stability, and calibration
  • Evaluating how predicted default probabilities align with realised outcomes over time
  • Tracking shifts in feature importance to identify changing drivers of risk

The models behaved differently as conditions changed. Random forests achieved stronger performance during relatively stable periods, but their AUC and calibration deteriorated more sharply as the data moved further from the conditions on which they were trained. Logistic regression achieved lower peak performance, but remained more stable over time.

Feature importance also changed across periods, particularly for variables such as loan grade and interest rate. This suggests that the relationships used by a model to assess risk are not fixed, and that strong historical performance does not necessarily imply reliable future behaviour.

The project highlights the importance of evaluating models in the context in which they will actually be used. When underlying relationships change, robustness and calibration can matter more than maximising predictive performance on historical data.

Quantum Computing in Bioinformatics (Internship)

Qiskit | Optimisation | Scientific Computing | Research

Completed a research internship at the Walter and Eliza Hall Institute of Medical Research, focusing on improving the usability and correctness of a quantum protein-folding model implemented in Qiskit. The work addressed technical and structural issues in a codebase handed down across multiple interns, where poor documentation and incorrect parameterisation limited practical use of the model.

A major part of the project involved improving the onboarding process. An extensive collection of academic readings was replaced with a structured glossary and more focused documentation, allowing new contributors to understand the key concepts required to work with the model. The implementation was also refactored with detailed inline commentary explaining theoretical background, modelling assumptions, and execution flow.

Key components included:

  • Designing a simplified onboarding pathway by consolidating essential concepts into a concise reference document
  • Refactoring the Jupyter notebook with structured inline documentation to improve readability and accessibility
  • Diagnosing failures in the Hamiltonian formulation caused by incorrectly scaled constraint penalties
  • Rescaling constraint terms to restore valid interaction interaction between amino acids
  • Replacing gradient-based optimisation with a non-derivative method better suited to the discrete and non-linear structure of the problem space
  • Testing the corrected model on simulated configurations to confirm functional behaviour while identifying remaining computational inefficiencies
  • Migrating the project from a Google Drive notebook to a structured GitHub repository, introducing version control and improving reproducibility
  • Producing technical report and handover documentation for future contributors

The project combined model debugging, optimisation, and research software improvement, with a strong emphasis on making technically complex work more understandable and maintainable for future contributors.

Where2 (Co-Founder & Developer)

React Native | Distributed Systems | Real-Time State Management

Co-founded and developed Where2, a mobile application for group coordination through shared itineraries, messaging, location discovery, and collaborative planning. The project began through the Melbourne Entrepreneurial Centre startup program and progressed from an early concept to its first public launch.

Worked primarily on the technical side of the product, including backend design, database structure, application logic, and real-time functionality. The system supported shared plans, location data, user accounts, access control, and synchronised updates across multiple users.

One of the most significant lessons from the project came from the architecture itself. As the application grew, the original implementation became increasingly difficult to maintain and extend. Much of the system was eventually rebuilt around a clearer service layer between the application and the database. This reduced direct coupling to the underlying data store and made later changes, including database migration and backend restructuring, considerably easier.

Key components included:

  • Designing relational data models for users, plans, locations, and shared application state
  • Implementing authentication, access control, and real-time synchronisation across multiple users
  • Contributing to a major architectural rebuild after limitations in the original implementation became difficult to manage
  • Introducing a service layer to separate application logic from database-specific implementation
  • Supporting the product through its first launch and subsequent reliability-focused development
  • Developing front-end functionality in React Native from iterative Figma designs

Toward the end of the project, work also focused on a place-field reliability system for evaluating competing claims about location data such as phone numbers, websites, descriptions, and tags. The design treated each source as an observation, combined corroborating evidence into confidence scores, and separated raw observations from derived confidence values and live write-back behaviour.

The project reinforced the importance of software architecture as a practical constraint on development. The most difficult problems were rarely isolated implementation issues. They tended to emerge from how components were structured, how tightly they depended on one another, and how difficult those dependencies made later change.

CSIRO Image2Biomass Datathon

Python | Deep Learning | Prediction Under Limited Data

Developed an end-to-end prediction pipeline to estimate pasture biomass from high-resolution aerial imagery as part of the CSIRO Image2Biomass datathon. The task involved translating large, unstructured visual data into quantitative estimates across multiple targets, including green, dry, and dead plant matter.

The primary challenge was working under limited labelled data and high variability in visual input. Images differed in scale, lighting, and spatial composition, requiring careful preprocessing to preserve meaningful structure while ensuring compatibility with model inputs.

A preprocessing pipeline was implemented to convert wide-format aerial images into consistent square inputs, applying normalisation and standardisation to stabilise feature extraction. A self-supervised Vision Transformer (DINO) was used as a feature extractor, with representations adapted for multi-output regression using PyTorch and optimised via mean squared error loss. This approach leveraged pre-trained representations to improve performance under constrained data conditions.

Key components included:

  • Designing a preprocessing pipeline to preserve spatial and textural information relevant to biomass estimation
  • Training and evaluating regression models across multiple data splits to assess generalisation
  • Combining model outputs using weighted averaging across data splits to improve robustness and reduce overfitting
  • Applying post-processing adjustments to correct systematic prediction bias

The project emphasised building a reliable prediction system under real-world constraints. Model performance depended not only on architecture, but on handling noisy inputs, limited labels, and variability in the data distribution. The final pipeline prioritised consistency and robustness over maximising performance on individual splits.

Flood Simulation

QGIS | Spatial Modelling | Risk Analysis

Developed a geospatial flood modelling framework to analyse inundation patterns and evaluate mitigation strategies across flood-prone regions, including areas in Queensland and St Kilda. The project treats flooding as a dynamic system influenced by terrain, water flow, and environmental conditions, rather than a static event.

Raster-based models were used to simulate flood propagation under varying severity scenarios, with interventions such as levees, temporary barriers, and sandbagging incorporated as constraints within the system. This allowed comparison of how different mitigation strategies alter flood extent, depth, and infrastructure exposure.

Key components included:

  • Constructing spatial models to simulate inundation under different environmental conditions
  • Incorporating intervention scenarios by modifying terrain and flow constraints
  • Analysing impact in terms of affected regions, infrastructure exposure, and changes in flood extent
  • Conducting long-term cost analysis over a 50-year horizon, incorporating durability, labour, and damage estimates

The project integrates spatial modelling with economic reasoning to support decision-making under uncertainty. Results highlight that optimal mitigation strategies depend not only on initial effectiveness, but on long-term resilience under changing conditions.

Decision-Making & Communication

Academic Misconduct Committee

University Governance | Evidence Evaluation | Institutional Decision-Making

Served as a student representative on the University of Melbourne Academic Misconduct Committee, participating in hearings relating to plagiarism, contract cheating, unauthorised collaboration, and misuse of AI tools.

The role involved reviewing sensitive and confidential case material, identifying inconsistencies in explanations, evaluating procedural fairness, and contributing to decisions with significant academic and disciplinary consequences. Cases often involved incomplete or conflicting information, requiring careful judgement under uncertainty rather than simple rule application.

Working within a confidential disciplinary process required maintaining discretion while handling sensitive personal information and evidence. The experience reinforced the importance of evidence-based reasoning, consistency of judgement, and understanding how individual behaviour interacts with institutional systems, incentives, and accountability structures.

Data Science Student Society & Responsible AI Development

Operational Coordination | AI Governance | Communication

Contributed to multiple student-led initiatives focused on industry engagement, AI systems, and responsible technology use through the Data Science Students Society (DSCubed) and Responsible AI Development.

Played a leading role in coordinating Industry Networking Night, managing communication between venue staff, student organisations, and external participants to organise logistics including catering, security, AV systems, seating, lighting, and event layout.

Also contributed to the DSCubed AI Projects Team, where work focused on improving AI-assisted onboarding and hiring processes used internally by the organisation. This included thinking about how AI systems could support operational workflows while remaining understandable and reliable for users.

Through Responsible AI Development and Green Impact initiatives, participated in discussions and evaluations surrounding responsible AI use, confidentiality risks, energy consumption, and broader social impacts of large-scale AI deployment.

Elevate Education

Adaptive Communication | Audience Engagement | Live Presentation Delivery

Delivered academic workshops and study-skills seminars to high school students across metropolitan and regional Victoria through Elevate Education. The role focused not only on presenting information, but on maintaining engagement and adapting delivery style dynamically across different classroom environments.

Presentations required continual adjustment based on audience behaviour, energy levels, and classroom dynamics. Sessions often involved balancing humour, authority, pacing, and audience participation in real time, particularly in environments where students were disengaged, fatigued, or difficult to involve. Delivery style, tone, and interaction patterns had to be adapted continuously depending on how different groups responded.

The role also required careful control of presentation mechanics such as voice projection, timing, room awareness, and conversational flow, particularly across classrooms with different acoustics, layouts, and group behaviour. Considerable preparation went into refining examples, delivery structure, and transitions to ensure material remained engaging, clear, and relatable to students from different backgrounds.

The experience reinforced the importance of behavioural awareness, adaptive communication, and maintaining clarity and engagement under unpredictable real-world conditions.

Skills

  • Systems & Infrastructure Python, C, SQL, React Native, Supabase, distributed systems, synchronisation workflows
  • Statistical Modelling & Machine Learning pandas, NumPy, scikit-learn, PyTorch, statsmodels, calibration and time-aware evaluation
  • Risk, Fraud & Behavioural Systems Signal reliability, anomaly detection, transaction analysis, network and behavioural analysis
  • Visualisation & Communication Tableau, Power BI, matplotlib, analytical writing, adaptive presentation delivery

Contact

Interested in roles across data science, technology risk, governance, and analytical decision-making, particularly where complex systems need to be understood carefully and evidence must be interpreted in context.