Skip to content
40 Reps

Study log · Nov to Dec 2022 · re-checked Oct 2026

40 reps, re‑checked.

At the end of 2022 I worked through 40 of Aman Kharwal's data science tutorials, one notebook each, to get my hands used to the everyday tools.

This site is the honest log of that run. It shows which notebooks actually ran, what they printed, and what changes when I re-check the results against simple baselines and proper test splits.

The log

one tile per notebook

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06
  7. 07
  8. 08
  9. 09
  10. 10
  11. 11
  12. 12
  13. 13
  14. 14
  15. 15
  16. 16
  17. 17
  18. 18
  19. 19
  20. 20
  21. 21
  22. 22
  23. 23
  24. 24
  25. 25
  26. 26
  27. 27
  28. 28
  29. 29
  30. 30
  31. 31
  32. 32
  33. 33
  34. 34
  35. 35
  36. 36
  37. 37
  38. 38
  39. 39
  40. 40
  • 33 ran
  • 3 partial
  • 3 did not run
  • 1 empty

40

notebooks, one per tutorial

33

ran from top to bottom

7

stopped early or never started

3 of 8

audits where the conclusion flips

Try it in your browser

Five of the audits come with a small tool you can play with. Everything runs in this page, and nothing you type or slide is sent anywhere.

What the re-check found

Eight deep-dives, each recomputed from licence-clear data or worked out from what the 2022 notebooks printed. Some conclusions hold. Several do not.
  • 01 · 05 · 09 · 28 · 32

    Conclusion flips

    Stock forecasts against 'tomorrow equals today'

    On a time-ordered split the Apple LSTM misses the close by 38.7% on average, against 1.7% for 'tomorrow equals today'. The linear model in notebook 05 also loses to 'no change' (0.50% versus 0.36%).

    Read the audit

  • 27

    InteractiveConclusion flips

    Four classifiers scored on their own training rows

    The tree's 0.98 implies 33 of 40 right on the unseen rows, about 82.5%. Logistic regression and the passive-aggressive model both scored exactly 257 of 400, which is what predicting 'no purchase' for everyone would score.

    Read the audit

  • 02

    InteractiveCan't tell from 2022

    99.97% accuracy on a 1-in-775 problem

    99.97% is 381 mistakes on 1,272,524 test rows. That is consistent with catching every fraud, or with missing almost a quarter of them, and the features include balances the dataset's author says must not be used.

    Read the audit

  • 03 · 14

    Weaker than it looked

    One lucky split: R² 0.906 from 40 rows

    Across 1,000 random splits of the book's Advertising data, the same model scores between 0.75 and 0.96. Repeated cross-validation settles at 0.886.

    Read the audit

  • 23 · 15

    InteractiveWeaker than it looked

    Thirty flowers cannot tell 93% from 100%

    The 1-nearest-neighbour model got 30 of 30 test flowers right, but the 95% interval runs from 88.4% to 100%. Over 1,000 other splits only 23% were perfect.

    Read the audit

  • 37

    InteractiveHolds up

    Spam detection holds up

    Yes. It scores 98.4% against 86.1% for 'always ham', catching 96% of spam with 93% precision. Ten-fold cross-validation agrees.

    Read the audit

How sure, and compared with what

Every audit number comes with an interval and every comparison with a paired test on the same items. The reasoning is written down, and one audit lets you put a language model against the 2022 model.

How the re-check works

  1. 01

    Read what actually ran

    A script reads the saved outputs of all 40 notebooks. It counts which cells ran, which failed and why, and pulls out the numbers each notebook printed. Nothing is re-executed for the log.

  2. 02

    Re-run with a baseline

    For eight notebooks, small Python scripts re-run the same method on data I am allowed to use, then add what was missing: a naive baseline, a time-ordered split, cross-validation, an interval and a paired test.

  3. 03

    Publish only derived numbers

    No tweets, posts, chat logs or raw datasets appear here, apart from a small attributed sample of a public SMS corpus for the optional LLM comparison. The site shows metrics, error curves and counts, and links each tutorial back to its author.

The tutorials, code style and narrative belong to Aman Kharwal's series. The descriptions, run log and audits are mine. More about this project, and how every number is tested on the methods page.