Skip to content
40 Reps

About

A study log, honestly kept

Personal project, 2022, revisited in 2026. Forty small exercises I did to build fluency, and a second look at what they really showed.

In November and December 2022 I worked through 40 projects from Aman Kharwal's series of 180 data science and machine learning projects with Python, published on thecleverprogrammer.com. The aim was repetition. Each day I loaded a dataset, cleaned and explored it, fitted something and read the output, then did it all again the next day. I followed each tutorial closely and saved one notebook per project.

Coming back to the notebooks in 2026, two things stood out. Several had not actually finished, which my old README did not say. And many of the headline numbers were measured in ways that flatter the model. Some notebooks shuffled a time series before splitting it, one scored its models on their own training rows, most had no baseline to compare with, and the two Iris notebooks set aside only 30 flowers for testing. This site is my attempt to report both honestly. Of the 40 notebooks, 33 ran from top to bottom, 3 stopped part-way, 3 failed at the start and 1 was never written.

None of this is a criticism of the tutorials, which do exactly what they promise: a short, runnable introduction to one idea at a time. The gaps are in how I used them, and they are the lessons I would most want a beginner to learn early.

Credit where it belongs

The project ideas, the code patterns and the narrative of each notebook come from Aman Kharwal's tutorials, and every project page links to the original. Because the notebooks follow those tutorials so closely, I do not republish their code or text here. The source repository stays private for the same reason, and because it holds third-party datasets that I am not allowed to redistribute.

The descriptions on this site, the run log, the notes on what I would do differently and every audit are mine. If you want to learn this material, read the tutorials themselves.

Provenance

The original notebooks are preserved unchanged in the repository's original/ folder, moved there with their git history. The run log on this site is extracted from their saved outputs by a script, so it shows what happened in 2022 rather than what happens if they are run again today.

The audits are new work. Small Python scripts run with uv re-create each result on data I can use, add the missing baseline or split, and write a few kilobytes of derived numbers that this site reads. Parity tests check the TypeScript code here against the Python libraries it replaces, and against numbers printed by the 2022 notebooks.

Data and licences

No tweet, Reddit post, chat log or user handle is served by this site, and no raw dataset except one: 400 masked messages from the CC BY 4.0 SMS Spam Collection, attributed, for the optional LLM comparison. Everything else is aggregates and metrics. Data cards.

Data used by each part of the site
PartSourceTerms and what is shown
Run log (all 40 notebooks)SourceSaved outputs of my own 2022 notebooksTerms and what is shownDerived numbers only
Stock forecasts (01, 05, 09, 28, 32)SourceYahoo Finance via yfinance, plus the CSVs bundled with the notebooksTerms and what is shownError metrics only, no price tables
Spam (37)SourceSMS Spam Collection, UCI Machine Learning RepositoryTerms and what is shownCC BY 4.0. Metrics, token counts, and a 400-message masked sample for the LLM comparison
Sales (03, 14)SourceAdvertising data, An Introduction to Statistical LearningTerms and what is shownDerived metrics only
Water quality (24)SourceWater Quality by Aditya Kadiwal, KaggleTerms and what is shownCC0
Iris (15, 23)SourceIris, UCI Machine Learning RepositoryTerms and what is shownCC BY 4.0
Car prices (36)SourceAutomobile, UCI Machine Learning RepositoryTerms and what is shownCC BY 4.0
Fraud (02)SourceNumbers printed by the 2022 notebookTerms and what is shownNothing re-run
Classifiers (27)SourceNumbers printed by the 2022 notebook, checked against a locally cached copy of its tableTerms and what is shownCounts only, table never published

Then and now

2022

  • Jupyter notebooks in conda environments on Linux and Windows
  • pandas and scikit-learn for most models
  • TensorFlow and Keras for the neural networks
  • NLTK's VADER for sentiment, AutoTS and LightGBM for forecasts and orders
  • Plotly, Matplotlib and seaborn for charts

2026

  • Next.js and TypeScript, styled with Tailwind CSS and a few shadcn/ui parts
  • Hand-drawn SVG charts with table views
  • Naive Bayes, confidence intervals and threshold maths ported to TypeScript, running in your browser
  • Vitest parity tests against scikit-learn, SciPy, statsmodels and R
  • Wilson, bootstrap and corrected cross-validation intervals, McNemar and Diebold-Mariano tests on every audit
  • An optional LLM comparison on your own API key, with an audit log kept in your browser
  • Audit scripts in Python, run with uv

The site needs no account and works fully without an API key. If you add one for the optional LLM comparison, it stays in your browser, and so does the audit log of every call. Start with the 40 notebooks.