I downloaded about 3,400 trading days of Apple prices and trained a two-layer LSTM to output the closing price from the same day's open, high, low and volume. The rows were shuffled before the 80/20 split.
The log
All 40 notebooks
Each card says what the notebook did in my words, how far it got in 2022, and the main number it printed. Filter by theme or status, or search by title, tutorial or number.
- 01Time series
- 02Classification
- 03Regression
- 04Regression
- 05Time series
- 06Text & sentiment
- 07Time series
- 08Text & sentiment
- 09Time series
- 10Time series
- 11Neural nets & vision
- 12Text & sentiment
- 13Neural nets & vision
- 14Regression
- 15Classification
- 16Regression
- 17Regression
- 18Text & sentiment
- 19Text & sentiment
- 20Text & sentiment
- 21Text & sentiment
- 22Text & sentiment
- 23Classification
- 24Classification
- 25Text & sentiment
- 26Text & sentiment
- 27Classification
- 28Time series
- 29Regression
- 30Exploration
- 31Regression
- 32Time series
- 33Classification
- 34Exploration
- 35Exploration
- 36Regression
- 37Text & sentiment
- 38Neural nets & vision
- 39Text & sentiment
- 40Exploration
- 33 ran
- 3 partial
- 3 did not run
- 1 empty
- Time series
- Regression
- Classification
- Text & sentiment
- Neural nets & vision
- Exploration
One tile per notebook, in the order I worked through them. Hover or focus a tile to see the number it printed, or open it for the full run record.
Notebooks
Showing 40 of 40
Ran Ran I encoded the transaction type as a number and trained a decision tree on type, amount and the sender's balance before and after the payment, then scored it on a shuffled 20% of the rows.
Ran I plotted sales against TV, radio and newspaper budgets, then fitted a linear regression on all three and scored it on one 80/20 split.
Empty I opened the notebook and wrote the introduction, then moved on to the next project without writing any code.
Ran I standardised the closing price, used today's value to predict the close five rows later with linear regression, and shuffled the rows before an 80/20 split.
Ran I lower-cased 2,304 product reviews, stripped punctuation and stop words and stemmed what was left. Then I scored each review with VADER, summed the positive, negative and neutral scores, and called the largest sum the overall mood.
Partial I downloaded 701 days of Bitcoin prices and started an AutoTS search for a 30-day forecast. I stopped the search after it had tried 137 models, so the forecast cell never ran.
Ran I turned 34 short articles into TF-IDF vectors, computed cosine similarity between every pair, and listed the four most similar titles for each article.
Partial The same set-up as project 01, applied to Netflix. The day's open, high, low and volume went in, the same day's close came out, and the rows were shuffled before an 80/20 split.
Ran I downloaded two years of Apple prices and drew five interactive charts: a line, a candlestick, a bar chart, a zoomed date range and a candlestick with range buttons.
Ran I flattened 28 by 28 Fashion-MNIST images and trained a 300-100-10 dense network with SGD for 30 epochs, holding out 5,000 training images for validation.
Ran I cleaned 2,838 posts, built a bag-of-words matrix with 7,842 terms, trained a Bernoulli naive Bayes classifier and tried it on two sentences I typed.
Ran I rebuilt an AlexNet-style network layer by layer in Keras for 227 by 227 colour images, compiled it, printed the summary and drew it with visualkeras. It was never trained.
Ran I fitted a one-variable linear regression of sales on TV spend and drew the fitted line over the scatter plot.
Ran I trained a 1-nearest-neighbour classifier on Iris and showed how to pass one new flower's four measurements to it as a NumPy array.
Ran I trained a decision tree to predict units sold from the total and base price across 150,150 store-week rows, then predicted demand for one price pair.
Ran I converted text columns to numbers, dropped rows with gaps and trained a random forest on ten wind, load and price columns to predict the half-hourly price.
Ran I turned 22,000 sentences in 22 languages into word counts and trained multinomial naive Bayes, which scored 95.3% on a 20% test split. A Hindi sentence I typed came back as Hindi.
Ran I scored 20,491 hotel reviews with VADER and labelled them positive, negative or neutral from the compound score with cut-offs at plus and minus 0.05, then saved the labelled table.
Ran I dropped every tweet with any missing field, cleaned the text, drew word clouds and summed VADER scores across the remaining 4,749 tweets.
Ran I trained multinomial naive Bayes on the headlines of 2,225 BBC articles in five categories and classified one headline I typed about an iPhone as tech.
Ran The same pipeline as project 20, run on 17,046 tweets about the Omicron variant. I dropped incomplete rows, cleaned the text, drew word clouds and summed the VADER scores.
Ran I explored the 150 flowers with a scatter plot, trained a 1-nearest-neighbour classifier on 120 of them and predicted the species of one new flower.
Ran I loaded 3,276 water samples, dropped every row with a missing value and plotted how many were safe to drink. The notebook stops there.
Ran I cleaned 24,783 tweets, scored them with VADER and summed the scores. Neutral came out on top.
Ran I loaded 80,019 tweets, kept the 54,267 with no missing fields, cleaned them, drew a word cloud and summed VADER scores.
Ran I trained KNN, a decision tree, logistic regression and a passive-aggressive classifier on age and salary to predict a purchase, then scored each one on all 400 rows.
Ran I loaded a year of Tata Motors ADR prices and let AutoTS search for the best five-day forecast, which it finished, producing a gentle upward path.
Ran I mapped sex and smoker to 0 and 1, checked correlations (smoking has by far the strongest link to charges) and trained a random forest on age, sex, BMI and smoker.
Ran I merged a movie table with about 900,000 ratings, drew the share of each rating value and listed the ten films with the most ten-out-of-ten ratings.
Ran I encoded store type, location type, holiday and discount as numbers and trained LightGBM to predict daily orders across 188,340 store-days.
Partial The same AutoTS set-up as project 28 on a year of Apple prices. The search was still running after 296 model attempts when I stopped it, so no forecast was printed.
Ran I explored who bought travel insurance by age, employment type and income, then trained a decision tree on seven columns and scored 81.4% on a 20% test split.
Ran I cleaned net worth from text to numbers, dropped rows with no age and drew charts of the richest ten plus the top five sources of wealth, industries and countries.
Ran I renamed the columns of a state-by-month table, drew a correlation heatmap, histograms by region and a sunburst of unemployment rates by region and state.
Ran I kept 14 numeric columns, trained a decision tree regressor on 80% of 205 cars and reported a score of 1.0.
- 37Ran
SMS spam
I turned 5,572 text messages into word counts, trained multinomial naive Bayes on two thirds of them, and checked one message I typed, which came back as spam.
Did not run The plan was to run a pre-trained YOLO detector through cvlib on a street photo and count the cars it found.
Did not run The plan was to parse an exported WhatsApp chat into date, time, author and message, then score each message with VADER.
Did not run The plan was to parse an exported chat and count messages, media, emojis and links per person, with word clouds for each author.