Taste Science

How Embedding-Based Recommendations Read Your Film Taste

AI movie recommendations based on taste beat collaborative filtering by representing films as vectors. Here's how the math captures what genre tags miss.

When Netflix shows you a row of films "because you watched Drive," the system isn't making a claim about Drive — it's making a claim about other people who watched Drive. The recommendation is really a statement: enough people whose watch history resembles yours also watched these films, so statistically, you might too. This is collaborative filtering, and it's been the dominant approach to recommendation for two decades. It works reasonably well at scale.

But it has a known failure mode, and if you have specific taste, you've felt it. The system keeps surfacing the same formal surface — action films with male protagonists, or whatever the genre bucket is — while completely missing the actual quality you loved about Drive: the sustained tension of its silences, the way Nicolas Winding Refn stages action as a kind of expressionist horror, the Ryan Gosling performance built almost entirely from withholding. None of that is encoded in the collaborative signal. What's encoded is the viewing behavior of several million people who also happened to watch it.

Embedding-based recommendations work from a different premise. Instead of what other people did, they try to encode what the film is.

A Geometric Intuition

Imagine mapping every film you've ever seen as a point in space. Not two-dimensional space — a space with many, many dimensions, each one representing some latent quality: emotional register, narrative pace, visual grammar, tonal temperature, thematic weight. In this space, films that feel similar are geometrically close, regardless of their genre labels or the popular audience they attracted.

Weekend (Godard, 1967) and Beau Is Afraid (Aster, 2023) have almost nothing in common by traditional taxonomy — different decades, national contexts, genre conventions, intended audiences. But in a well-trained embedding space, they might end up relatively close: both are structurally hostile, both use dread as a sustained atmospheric condition rather than a punctuating shock, both treat bourgeois comfort as a site of horror. Someone who loves one and is open to discovering the other has made a coherent move.

This is what embeddings capture that genre tags don't: texture, register, feel. The things that actually drive why you love a film.

How the Representation Gets Built

An embedding is a list of numbers — a vector — that represents an entity in this multi-dimensional space. Generating good film embeddings requires some form of representation learning: exposing a model to a large body of information about films (critical writing, audience responses, plot descriptions, visual and audio features for those models that can handle them) and training it to place films that share latent qualities near each other.

The training signal matters enormously. Train primarily on plot summaries and you'll get embeddings that cluster films by narrative structure and miss everything about tone. Include substantial critical discourse — the vocabulary that film writing has developed over a century to describe how films work — and you get richer representations. The embedding inherits some of the expressiveness of that critical tradition.

A viewer also gets an embedding — derived from the films they've watched and how they've rated them. If you've consistently rated slow-burn, long-take, formally rigorous films highly and marked fast-cut genre product down, your embedding will reflect that. You'll sit in a part of the space populated by films like yours.

The recommendation problem then becomes: find films near your position that you haven't seen.

Why Two Tonally Similar Films from Different Genres End Up Neighbors

Consider Annihilation (Garland, 2018) and Picnic at Hanging Rock (Weir, 1975). Standard collaborative filtering would rarely connect them — different audiences, different eras, wildly different genre classifications. But critical writing often describes them in similar terms: the uncanny encroachment of nature as a force that erases human individuality, dread that never resolves into explanation, female protagonists confronting something the film refuses to rationalize. A model trained on substantial film criticism would push these two films closer together in embedding space than, say, Annihilation and another 2018 science fiction film that shares none of its tonal qualities.

This is the core of what "taste-based" means as opposed to "genre-based" or "popularity-based." The representation is of how the film works on you — its phenomenology, loosely speaking — rather than what shelf it would sit on in a video shop.

The Honest Limits

Embeddings aren't magic, and being honest about their limits matters.

They still depend heavily on training data. Films with thin critical coverage — much of world cinema outside the European-American canon, recent films from national traditions that receive less English-language critical attention, older films whose reception history has faded — will have noisier embeddings or may cluster incorrectly. A system trained primarily on English-language film writing will have a latent bias toward the critical frameworks that writing has developed.

They also can't escape the cold-start problem entirely. If you've only logged a handful of films, your user embedding is underspecified — there isn't enough signal to place you precisely. The recommendations will be reasonable, but the real refinement happens over time, as the viewing record grows and the system has more to work with.

And embedding similarity isn't the same as recommendation quality. A film that's geometrically very close to your taste profile might still be one you find merely adequate because it's doing the same thing you've already seen done better elsewhere. Novelty and surprise have to be balanced against similarity.

How This Differs from What Netflix Is Doing

Netflix's recommendation engine is heavily collaborative — it operates at a scale where aggregated viewing behavior is the most powerful signal available. They have hundreds of millions of viewers, and the correlations in that data are genuinely useful, particularly for predicting whether a viewer will start a film (which is what their metrics optimize for).

Intertitle is built for a different goal: connecting a specific viewer, with specific taste, to films they'll find genuinely meaningful — not just films they'll click on. The corpus is different (a curated film database, not a streaming catalog optimized for engagement), the signal is different (Letterboxd ratings reflect considered opinion, not passive viewing behavior), and the optimization target is different (taste match, not predicted engagement).

The embedding approach means that when Intertitle recommends something, it's making a claim about the film — about what kind of film it is — mapped against a claim about you. If it's wrong, the error is interesting: it means the embedding misread either the film or your taste, which is a more useful failure mode than "other people also watched this."

Your Letterboxd history is the training data for your own embedding. The more complete it is, the more precisely the system can read you. Intertitle builds that representation from what you've already logged, then uses it to find what you haven't seen yet.

Intertitle

Intertitle reads your taste
and tells you what to watch tonight.

Import your Letterboxd export. One recommendation, matched to your actual watching history — not the algorithm's guess.

↓ Download on iOS

Free · no credit card