Skip to content
40 Reps

The log

All 40 notebooks

Each card says what the notebook did in my words, how far it got in 2022, and the main number it printed. Filter by theme or status, or search by title, tutorial or number.

  1. 01
  2. 02
  3. 03
  4. 04
  5. 05
  6. 06
  7. 07
  8. 08
  9. 09
  10. 10
  11. 11
  12. 12
  13. 13
  14. 14
  15. 15
  16. 16
  17. 17
  18. 18
  19. 19
  20. 20
  21. 21
  22. 22
  23. 23
  24. 24
  25. 25
  26. 26
  27. 27
  28. 28
  29. 29
  30. 30
  31. 31
  32. 32
  33. 33
  34. 34
  35. 35
  36. 36
  37. 37
  38. 38
  39. 39
  40. 40
  • 33 ran
  • 3 partial
  • 3 did not run
  • 1 empty

Notebooks

Showing 40 of 40

  • Ran

    I downloaded about 3,400 trading days of Apple prices and trained a two-layer LSTM to output the closing price from the same day's open, high, low and volume. The rows were shuffled before the 80/20 split.

  • Ran

    I encoded the transaction type as a number and trained a decision tree on type, amount and the sender's balance before and after the payment, then scored it on a shuffled 20% of the rows.

  • Ran

    I plotted sales against TV, radio and newspaper budgets, then fitted a linear regression on all three and scored it on one 80/20 split.

  • Empty

    I opened the notebook and wrote the introduction, then moved on to the next project without writing any code.

    Regression

    No number printed in 2022

  • Ran

    I standardised the closing price, used today's value to predict the close five rows later with linear regression, and shuffled the rows before an 80/20 split.

    Time series

    Test R² (linear regression, 5 days ahead): 0.639

  • Ran

    I lower-cased 2,304 product reviews, stripped punctuation and stop words and stemmed what was left. Then I scored each review with VADER, summed the positive, negative and neutral scores, and called the largest sum the overall mood.

    Text & sentiment

    Overall verdict: Neutral

  • Partial

    I downloaded 701 days of Bitcoin prices and started an AutoTS search for a 30-day forecast. I stopped the search after it had tried 137 models, so the forecast cell never ran.

    Time series

    AutoTS models tried before the run stopped: 137

  • Ran

    I turned 34 short articles into TF-IDF vectors, computed cosine similarity between every pair, and listed the four most similar titles for each article.

  • Partial

    The same set-up as project 01, applied to Netflix. The day's open, high, low and volume went in, the same day's close came out, and the rows were shuffled before an 80/20 split.

  • Ran

    I downloaded two years of Apple prices and drew five interactive charts: a line, a candlestick, a bar chart, a zoomed date range and a candlestick with range buttons.

    Time series

    No number printed in 2022

  • Ran

    I flattened 28 by 28 Fashion-MNIST images and trained a 300-100-10 dense network with SGD for 30 epochs, holding out 5,000 training images for validation.

    Neural nets & vision

    Validation accuracy after epoch 30: 88.9%

  • Ran

    I cleaned 2,838 posts, built a bag-of-words matrix with 7,842 terms, trained a Bernoulli naive Bayes classifier and tried it on two sentences I typed.

    Text & sentiment

    Posts vectorised: 2,838

  • Ran

    I rebuilt an AlexNet-style network layer by layer in Keras for 227 by 227 colour images, compiled it, printed the summary and drew it with visualkeras. It was never trained.

    Neural nets & vision

    AlexNet parameters: 41,546,506

  • Ran

    I trained a 1-nearest-neighbour classifier on Iris and showed how to pass one new flower's four measurements to it as a NumPy array.

    Classification

    Prediction for the sample flower: Iris-setosa

  • Ran

    I trained a decision tree to predict units sold from the total and base price across 150,150 store-week rows, then predicted demand for one price pair.

    Regression

    Units predicted at the sample prices: 27.00

  • Ran

    I converted text columns to numbers, dropped rows with gaps and trained a random forest on ten wind, load and price columns to predict the half-hourly price.

    Regression

    Predicted SMPEP2 price for the sample row: 65.92

  • Ran

    I turned 22,000 sentences in 22 languages into word counts and trained multinomial naive Bayes, which scored 95.3% on a 20% test split. A Hindi sentence I typed came back as Hindi.

    Text & sentiment

    Test accuracy (multinomial naive Bayes): 95.3%

  • Ran

    I scored 20,491 hotel reviews with VADER and labelled them positive, negative or neutral from the compound score with cut-offs at plus and minus 0.05, then saved the labelled table.

    Text & sentiment

    Reviews labelled positive: 18,831

  • Ran

    I dropped every tweet with any missing field, cleaned the text, drew word clouds and summed VADER scores across the remaining 4,749 tweets.

    Text & sentiment

    Overall verdict: Neutral

  • Ran

    I trained multinomial naive Bayes on the headlines of 2,225 BBC articles in five categories and classified one headline I typed about an iPhone as tech.

    Text & sentiment

    Category for one typed headline: tech

  • Ran

    The same pipeline as project 20, run on 17,046 tweets about the Omicron variant. I dropped incomplete rows, cleaned the text, drew word clouds and summed the VADER scores.

    Text & sentiment

    Overall verdict: Neutral

  • Ran

    I explored the 150 flowers with a scatter plot, trained a 1-nearest-neighbour classifier on 120 of them and predicted the species of one new flower.

    Classification

    Prediction for the sample flower: Iris-setosa

  • Ran

    I loaded 3,276 water samples, dropped every row with a missing value and plotted how many were safe to drink. The notebook stops there.

  • Ran

    I cleaned 24,783 tweets, scored them with VADER and summed the scores. Neutral came out on top.

    Text & sentiment

    Overall verdict: Neutral

  • Ran

    I loaded 80,019 tweets, kept the 54,267 with no missing fields, cleaned them, drew a word cloud and summed VADER scores.

    Text & sentiment

    Overall verdict: Neutral

  • Ran

    I trained KNN, a decision tree, logistic regression and a passive-aggressive classifier on age and salary to predict a purchase, then scored each one on all 400 rows.

  • Ran

    I mapped sex and smoker to 0 and 1, checked correlations (smoking has by far the strongest link to charges) and trained a random forest on age, sex, BMI and smoker.

    Regression

    Correlation of smoker with charges: 0.79

  • Ran

    I merged a movie table with about 900,000 ratings, drew the share of each rating value and listed the ten films with the most ten-out-of-ten ratings.

    Exploration

    Ten-star ratings for the most-loved title: 1,479

  • Ran

    I encoded store type, location type, holiday and discount as numbers and trained LightGBM to predict daily orders across 188,340 store-days.

    Regression

    Test rows predicted (LightGBM): 37,668

  • Partial

    The same AutoTS set-up as project 28 on a year of Apple prices. The search was still running after 296 model attempts when I stopped it, so no forecast was printed.

    Time series

    AutoTS models tried before the run stopped: 296

  • Ran

    I explored who bought travel insurance by age, employment type and income, then trained a decision tree on seven columns and scored 81.4% on a 20% test split.

    Classification

    Test accuracy (decision tree): 81.4%

  • Ran

    I cleaned net worth from text to numbers, dropped rows with no age and drew charts of the richest ten plus the top five sources of wealth, industries and countries.

    Exploration

    Billionaires listed: 2,755

  • Ran

    I renamed the columns of a state-by-month table, drew a correlation heatmap, histograms by region and a sunburst of unemployment rates by region and state.

    Exploration

    Mean estimated unemployment rate (%): 12.24

  • Ran

    I kept 14 numeric columns, trained a decision tree regressor on 80% of 205 cars and reported a score of 1.0.

    Regression

    Reported R² (scored against its own predictions): 1.000

  • Ran

    I turned 5,572 text messages into word counts, trained multinomial naive Bayes on two thirds of them, and checked one message I typed, which came back as spam.

  • Did not run

    The plan was to run a pre-trained YOLO detector through cvlib on a street photo and count the cars it found.

    Neural nets & vision

    No number printed in 2022

  • Did not run

    The plan was to parse an exported WhatsApp chat into date, time, author and message, then score each message with VADER.

    Text & sentiment

    No number printed in 2022

  • Did not run

    The plan was to parse an exported chat and count messages, media, emojis and links per person, with word clouds for each author.

    Exploration

    No number printed in 2022