Production
Working on data products and process analytics
At dmTECH I work on data products, process analytics, and the pipelines, internal packages, and deployments around them across Python, PySpark, Docker, Terraform, and Google Cloud.
Machine Learning / Applied AI / Data Science
Master's student at TU Berlin working on my thesis, graph neural networks and LLM-derived embeddings for ransomware detection. I build data products at dmTECH, and I'm exploring an early-stage idea around digital twins for critical infrastructure.
Trajectory
Production
At dmTECH I work on data products, process analytics, and the pipelines, internal packages, and deployments around them across Python, PySpark, Docker, Terraform, and Google Cloud.
Research
My current research explores how LLM-derived labels, embeddings, and related signals can enhance graph structures and graph learning workflows.
Teaching
I taught algorithms and programming at TU Berlin and learned how to explain technical concepts clearly and concretely.
Selected work
Project
LangChain-based LLM agent to autonomously generate, compile, and debug Java plugins for Google's Tsunami Security Scanner. Designed a two-stage workflow combining code generation with compiler feedback for self-correction, enabling fully automated plugin development for OWASP Juice Shop vulnerabilities. In evaluation, the agent produced 16 working vulnerability detection plugins, all without manual code fixes.
Project
Geospatial Computer Vision pipeline for automated building detection from multispectral satellite imagery. Integrated Sentinel-2 and OpenStreetMap data, generated segmentation masks, prepared stratified training/validation datasets, and trained PyTorch-based CNN and U-Net models with hyperparameter tuning and augmentation.
Project
Project to detect station stops of Berlin commuter trains using only phone sensor data that apps can access without special permissions (magnetometer, accelerometer, gravity, and gyroscope streams). The data is resampled and transformed into time-based features that are less sensitive to how the phone is oriented and used to train an ML model that predicts stop intervals for each ride. Uses HistGradientBoosting with 328 engineered features including lag and diff features. Achieved 0.722 Jaccard score and top leaderboard position.
Technical focus
I work most comfortably where machine learning, data engineering, and product-minded implementation meet. For a fuller picture, the CV page covers coursework, teaching, volunteering, and the broader toolset I worked with.
If the role needs someone comfortable moving between research context, implementation, and explanation, that is the space I like working in. I'm also open to early-stage collaborations.