HELLO, I’M

Caroline Kim

I build data and machine-learning systems end to end — scraping and cleaning real-world data, modeling it, and shipping the result as tested, containerized, cloud-deployed services.

Over one summer at Johns Hopkins (EP 605.256, Modern Software Concepts in Python) I took a single messy dataset — 100,000 graduate-admissions records — from raw scrape to a fine-tuned DistilBERT model deployed behind a Flask app, hardening every layer along the way: tests at 100% coverage, injection-safe SQL, Docker microservices, AWS deployment, and tracked ML experiments.

CORE SKILLS

The tools I reach for most, and what I’ve built with them.

FOCUS
Cloud & DevOps

Testing & CI, secure SQL, Docker microservices, AWS (S3, SageMaker, EC2).

Data Engineering

Scraping, cleaning with audit ledgers, PostgreSQL, idempotent pipelines.

ML & Modeling

Clustering, neural networks from scratch, LM fine-tuning, experiment tracking.