DEV Community

Ayush S Pangaonkar
Ayush S Pangaonkar

Posted on

A Book Recommender That Nails Harry Potter and Fumbles Seven Plays

Let me start with the good result. I asked my book recommender for books similar to Harry Potter and the Order of the Phoenix, and it returned:

Rank Recommended title Similarity
1 Harry Potter and the Chamber of Secrets 0.782
2 Harry Potter and the Sorcerer's Stone 0.778
3 Harry Potter and the Half-Blood Prince 0.768
4 Harry Potter and the Prisoner of Azkaban 0.765
5 Harry Potter Boxed Set, Books 1-5 0.764

Nice. Now the other one. I asked for books similar to Seven Plays:

Rank Recommended title Similarity
1 Buried Child 0.458
2 See You Around, Sam! 0.263
3 Sam Walton: Made In America 0.246

The second and third results are there because of the name "Sam", not because the books are alike. Same model, very different quality. Here is why.

How it works

The dataset has 11,127 books with bookID, title, authors and average_rating, and no missing values.

  1. I build a book_content field by joining each book's title and authors.
  2. TfidfVectorizer (English stopwords removed) turns that text into vectors. The matrix is 11,127 by 17,937, and only 0.04% of its cells are non-zero.
  3. linear_kernel computes cosine similarity between every pair of books, giving an 11,127 by 11,127 matrix, around 124 million scores.
  4. For a queried title, I return the top 10 most similar books, excluding the book itself.

A simplified sketch of the idea:

tfidf = TfidfVectorizer(stop_words="english")
matrix = tfidf.fit_transform(books["book_content"])
similarity = linear_kernel(matrix, matrix)
Enter fullscreen mode Exit fullscreen mode

Why Harry Potter works and Seven Plays doesn't

The only signal I gave the model is title and author text. A long series with a consistent author and naming pattern shares a lot of words, so it clusters well, with similarities between 0.66 and 0.78. A title with no such pattern leaves the model matching on incidental word overlap.

This is a content-based recommender, so it relies entirely on item text. There is no genre, no description, and no user ratings in it.

What I would do next

Add genre or description text to the content string, which gives the vectors something real to compare. If user rating histories were available, collaborative filtering could be layered on top.

The lesson I took away: a recommender is only as good as the features behind it, and testing the awkward queries tells you more than testing the easy ones.


Code: github.com/bluntjudg/Book-Recommendation-System-

Series: Part 2 of 7 in my ML fundamentals revisit. Next in the series: House Price Predictions.

Live projects I built after these basics:

More of my work is on GitHub.

Top comments (0)