Let me start with the good result. I asked my book recommender for books similar to Harry Potter and the Order of the Phoenix, and it returned:
| Rank | Recommended title | Similarity |
|---|---|---|
| 1 | Harry Potter and the Chamber of Secrets | 0.782 |
| 2 | Harry Potter and the Sorcerer's Stone | 0.778 |
| 3 | Harry Potter and the Half-Blood Prince | 0.768 |
| 4 | Harry Potter and the Prisoner of Azkaban | 0.765 |
| 5 | Harry Potter Boxed Set, Books 1-5 | 0.764 |
Nice. Now the other one. I asked for books similar to Seven Plays:
| Rank | Recommended title | Similarity |
|---|---|---|
| 1 | Buried Child | 0.458 |
| 2 | See You Around, Sam! | 0.263 |
| 3 | Sam Walton: Made In America | 0.246 |
The second and third results are there because of the name "Sam", not because the books are alike. Same model, very different quality. Here is why.
How it works
The dataset has 11,127 books with bookID, title, authors and average_rating, and no missing values.
- I build a
book_contentfield by joining each book's title and authors. -
TfidfVectorizer(English stopwords removed) turns that text into vectors. The matrix is 11,127 by 17,937, and only 0.04% of its cells are non-zero. -
linear_kernelcomputes cosine similarity between every pair of books, giving an 11,127 by 11,127 matrix, around 124 million scores. - For a queried title, I return the top 10 most similar books, excluding the book itself.
A simplified sketch of the idea:
tfidf = TfidfVectorizer(stop_words="english")
matrix = tfidf.fit_transform(books["book_content"])
similarity = linear_kernel(matrix, matrix)
Why Harry Potter works and Seven Plays doesn't
The only signal I gave the model is title and author text. A long series with a consistent author and naming pattern shares a lot of words, so it clusters well, with similarities between 0.66 and 0.78. A title with no such pattern leaves the model matching on incidental word overlap.
This is a content-based recommender, so it relies entirely on item text. There is no genre, no description, and no user ratings in it.
What I would do next
Add genre or description text to the content string, which gives the vectors something real to compare. If user rating histories were available, collaborative filtering could be layered on top.
The lesson I took away: a recommender is only as good as the features behind it, and testing the awkward queries tells you more than testing the easy ones.
Code: github.com/bluntjudg/Book-Recommendation-System-
Series: Part 2 of 7 in my ML fundamentals revisit. Next in the series: House Price Predictions.
Live projects I built after these basics:
- ATS Resume Analyzer, a Streamlit app that scores resumes against job descriptions
- AI-Based Loan Verification System, a Streamlit app for automated loan eligibility checks
- Subreach, a two-agent Reddit tool
More of my work is on GitHub.
Top comments (0)