skinDisease — Dermatology Classification
Deep learning that flags 23 skin conditions from a photo — and shows its work.

My Role
Sole developer. I built the full training pipeline — data preparation and stratified splits, a two-phase fine-tuning loop, evaluation with confusion matrices, and Grad-CAM explainability — with a config-driven codebase and unit tests so the whole thing is reproducible, not a one-off notebook.
Technology
Python · PyTorch · EfficientNet-B0 · timm · Albumentations · Grad-CAM · Mixed precision (AMP)
Measurable Impact
- Skin conditions classified23
- Training images (DermNet)~19,500
- Top-1 accuracy~75%*
- ExplainabilityGrad-CAM
* Pending verification
The Problem
Skin conditions are among the most common reasons people see a doctor, yet dermatologists cluster in big cities. A family in a rural district can wait weeks and travel hours for an answer to something a photo could triage. I wanted to see how far an honest, explainable model could go toward closing that gap — not to replace a doctor, but to help decide who needs one soon.
My Role
I built it alone. Every part — how the data is split, how the model is trained, how results are measured and how they are explained — was a decision I made and can defend, and I wrote it as a reproducible codebase with tests rather than a single notebook that only runs on my machine.
The Data Challenge
Medical datasets are brutally imbalanced: a few conditions have thousands of images, others only a handful. A model that ignores the rare classes can still look accurate on paper while being useless for the patients who need it most. I fought that with a weighted sampler so rare conditions appear every epoch, a class-weighted loss for a stronger gradient on them, and skin-aware augmentation (including elastic deformation) to squeeze more signal from limited images.
Model & Training
I chose EfficientNet-B0 for its accuracy-to-size ratio and fine-tuned it in two phases — freeze the pretrained backbone to warm up a new head, then unfreeze to adapt gently — with a warmup-plus-cosine learning-rate schedule and mixed-precision training for roughly a 2× speedup. To be sure the choice was principled, I benchmarked it against ResNet-50 and ConvNeXt rather than trusting one run.
Explainability & Ethics
A label alone is dangerous in medicine. Grad-CAM overlays where the model was looking, so a clinician can tell whether it focused on the lesion or on a shadow — and reject it when it's wrong. I am deliberate about the framing too: this is a screening aid to prioritise, never a diagnosis, and I keep the words 'aid' and 'diagnose' apart on purpose.
Honest Results
On this dataset EfficientNet-B0 reaches an expected top-1 around 75% and top-5 around 93%, edging out ResNet-50 — but I report these as expected ranges, not verified claims, because they shift with hardware and random seed and I have not run an independent clinical evaluation. Saying '~75%, pending verification' is less impressive than a bold number, and far more trustworthy — which, for a medical tool, is the only thing that counts.
Why It Matters & What's Next
The goal was never a leaderboard score; it was triage for people far from a specialist. Next, I want a proper held-out clinical evaluation with a doctor in the loop, probability calibration so the model knows when it's unsure, and a lightweight on-device version that works without a connection.
Technology
What I Learned
- In medical AI the rare classes are the whole point; average accuracy can hide exactly the cases that matter.
- An explanation you can overrule is worth more than a confident label you can't.
- Reporting a number honestly — as a range, pending verification — is a design choice, not a weakness.
What comes next
- Run a held-out clinical evaluation with a dermatologist and calibrate the model's confidence.
- Ship a lightweight offline version for clinics with poor connectivity.
Last updated: 2026-08-15