Projects

Selected work with public code. These are earlier team and course projects; my production experience lives under Experience.

Systems · Student team project · 2019

A live chart of trending artists

KafkaSpark StreamingCassandraDash
AI-generated concept illustration · not a dashboard
AI-generated concept illustration · not a dashboard
Problem
Turn a stream of tweets sharing Spotify tracks into a rolling chart of the artists people were sharing.
What we built
With a classmate, I built a Python pipeline: ingest links into Kafka, enrich tracks through Spotify, process with Spark Streaming, store counts in Cassandra, and redraw the top 20 in Dash.
Decision / limit
The demo used one copy of every component on a laptop. Rereading it exposed two quiet failures: restart offsets were not saved, and rows keyed by artist and second could overwrite separate shares.
Result / evidence
A working dashboard and recorded demo, not a production deployment. The later write-up separates what we built from the checkpointing, event-time and idempotency fixes I would make now.
Back to project

Data · KTH course team project · 2019

When a linear model beat the CNN

Pythonscikit-learnKerasNLP
AI-generated concept illustration · not model output
AI-generated concept illustration · not model output
Problem
Compare sentiment classifiers on tweets, then ask how much the result depends on the data rather than the model.
What we built
With two classmates, I compared a lexicon baseline, Naive Bayes, a linear SVM and a small CNN, including cleaning, sparse text features and experiments with stopword lists.
Decision / limit
Removing the standard stopword list also removed "not". A later review found a more serious limit: the CNN was overfitting, so its comparison with the SVM was not a clean test of model families.
Result / evidence
The saved Twitter test results put the linear SVM at 0.820 accuracy and the CNN at 0.777. These are historical course-project results, not a benchmark claim. The write-up also documents inconsistent AUC reporting and experiments that never actually ran.
Back to project

Data · KTH DD2424 course team project

Adding colour to grayscale images

TensorFlowKerasInception-ResNet-v2
Repo output · grayscale / prediction / original
Repo output · grayscale / prediction / original
Problem
Explore whether a neural network can predict colour from a grayscale image, combining local image detail with high-level features.
What we built
Our three-person team reimplemented a published colorization approach with modifications: an encoder, a pretrained feature extractor, a fusion step and a decoder that predicts colour channels.
Decision / limit
Before broader training, we overfit one image to check the network wiring. Project time limited the ImageNet experiment to 2,000 images, so this is a small reproduction, not a claim of general performance.
Result / evidence
When we scaled training with precomputed features stored in TFRecords, the model produced brownish images. We suspected a TFRecords implementation problem, but did not confirm the cause. Our smaller mixed-class model handled grass and sky better than clothing details. The public repository contains notebooks, the architecture, output examples and a report. It does not establish production use or a measured improvement over a baseline.
Back to project
Case study