Presenting a tag’s contents in a logical order—sorting by semantic similarity, which looks like a “sort by magic”:

But we can go further. Why is a mega-tag such a pain to read or search? Well, one problem is that they tend to be a giant messy pile with no order aside from reverse-chronological. Reverse-chronological order is bad in many cases, even blogs (consider a multi-part series where you can only reach them all by the tag, which of course shows you the series in the worst possible order!), and is used simply because… how else are you going to sort them? At least that shows you the newest ones, which is not always a good order, but at least is an order. To go beyond that, you’d need some sort of semantic understanding, the sort of deeper understanding that a human would have (and of course, human brains, particularly the brain of your human user, are too expensive to use to provide some sensible order).

Fortunately, we have those document embeddings at hand. We could try clustering with k-means & titling with an LLM again, and display each cluster one after another, treating the clusters as ‘temporary’ or ‘pseudo’ tags. (We can easily name the anonymous clusters with an LLM by feeding in the metadata like titles and asking for a tag name. The user may reject it, but even a wrong tag name is extremely helpful for making it obvious what the right one is, and breaks “the tyranny of the blank page” and decision fatigue.) Clusters don’t respect our 2D reading order, but there are alternate ways of clustering which are intended to project the high-dimensional embedding’s clustering geometry down to fewer dimensions, like 2D, for use in graphs, or even 1D—eg. t-SNE or UMAP.

I don’t know if they work well in 1D, but if they work better at slightly higher dimensionality, then it can be easily turned into a sequence with minimum total distance as a traveling salesman problem. A simple way to ‘sort’ which doesn’t require heavy-weight machinery is to ‘sort by semantics’: however, not by distance from a specific point, but greedily pairwise. One selects an arbitrary starting point (‘most recent item’ is a logical starting point for a tag), finds the ‘nearest’ point, adds it to the list, and then the nearest unused point to that point, and so on recursively.

I find with Gwern.net annotations, the greedy list sorting algorithm works surprisingly well. It naturally produces a fairly logical sequence with occasional ‘jumps’ as the latent cluster changes, in contrast to the naive ‘sort by distance’ which would tend to ‘ping-pong’ or ‘zig-zag’ back & forth across clusters based on slight differences in distance.

This would implicitly expose the underlying structure by preserving the local geometry (even if the ‘global’ shape doesn’t make sense), and help a reader skim through, as they feel ‘hot’ and ‘cold’, and can focus on the region of the tag which seems closest to what they want. (And if the tag really needs to be chronological, or the embedding linearization is bad, there can just be a setting to override that.)

This approach would work for anything that can be usefully embedded, and would probably work even better for images given how hard it is to organize images & how good image embeddings like CLIP have become (eg. Concept or SOOT). This will wind up inevitably producing some abrupt transitions between clusters, but that tells you where the natural categories are, and you can easily drag-and-drop the cluster of images into folders & redo the trick inside each directory. This would make it easy to drag-and-drop a set of datapoints, and select them, and define a new tag which applies to them.

And because embeddings are such widely-used tools, there are many tricks one can use. For example, the default embedding might not put enough weight on what you want, and might wind up clustering by something like ‘average color’ or ‘real-world location’. But embedders can be prompted to target specific use-cases, and if that is not possible, you can manipulate the embedding directly based on the embedding of a specific point such as a prototypical file (or a query/keyword prompt if text-only or using a cross-modal embedding like CLIP’s image+text): embed the new file or prompt, weighted multiply all the others by it (or something), then re-organize. (And you can of course finetune any embedding model with the user’s improvements, contrastively: push further apart the points that the user indicated were not as alike as the embedding indicated, and vice-versa. This works most easily with the model that generated the embedding, but one can come up with tricks to finetune other models as well on the user actions.)

Or you can experiment with “embedding arithmetic”: if the default 2D layout is unhelpful, because the most visible variation is focused on unhelpful parts, one can ‘subtract’ embeddings to change what shows up. And you can do this with any number of embeddings by doing arithmetic on them first. For example, you can ‘subtract’ a tag X from every datapoint to ignore their X-ness, by averaging every datapoint with tag X to get a “prototypical X”; the new embeddings are now “what those datapoints mean besides the concept encoded tag X”. (If the right tag or datapoint doesn’t exist which emphasises the right thing—just make one up!) By sequentially subtracting, one can look through the dataset as a whole for ‘missing’ tags; indeed, if every tag is subtracted, the residual clusters might still be surprisingly meaningful, because they have structure that no tag yet encoded. One could also try adding in order to emphasize a specific X.