-
Choose-Your-Own-Adventure generative fiction for efficiency/editing (2021-06-06)
-
Try directly optimizing reward generation (2019-12-16):
-
Progressive Growing for Autoregressive Models (2021-09-28):
autoregressive models like DALL·E 1 admit a simple and appealing curriculum by generating the sequences for each resolution, small to large, in order, enabling stabler faster learning, more global coherency, and allowing adaptive sampling (only generating full-scale as necessary).Implemented in April 2024 by Tian et al 2024, where “VAR” showed excellent scaling laws & performance.Progressive growing for images was introduced by ProGAN: it trained on 4px images, then added layers to the GAN to generate 16×16px images, then 32×32px and so on. This stabilized the GAN, and resulted in then-remarkably-fast learning. The GAN could learn the global appearance of an image and then the next finer level of detail, one level at a time, without the difficult learning problem of learning every aspect of an image at once.Autoregressive models like DALL·E 1 or PixelRNN long before it do not have any clear progressive growing. They start in the upper left corner, predict the first pixel, then predict the second pixel given the first pixel, and so on. For a GAN being trained on images, which just takes a random seed and spits out a finished image, it’s easy enough to see how progressive growing ought to work. But what does that mean for an autoregressive model? Approaches like hierarchical PixelCNNs or training on wavelet encodings (eg. JPEG-style Fourier coefficients, like Not-so-BigGAN) have either not been compelling or have been neglected.I suggest that autoregressive models be reformulated to predict images of a specific resolution conditional on the smaller resolutions. The AR model doesn’t just start predicting the pixels of a 1,024×1,024px image, it begins by predicting the pixels of the 16×16px image first (possibly based on a text description or a noise input); then, it predicts the 32×32px given the 16×16px image, and so on. (So as a sequence, it is simply the same image repeated at increasing resolution, as a “mipmap”, which adds only 33% overhead compared to the highest-resolution image. The prefix can also include ‘sketches’ like semantic segmentation, like Make-A-Scene does.)This provides the benefits of progressive growing during training: it starts by just predicting 16×16px images; when the prediction loss falls enough, it can then start predicting the 32px sequences (there is no need to make it predict the 16px ones now, and they can just be included in the prefix); then the 64px and so on. (So images which are “easy” automatically get skipped in a principled way, spending compute on either hard images, or hard resolutions of easy images.) It saves compute because one doesn’t train at the full final scale (remember, a 1024px image has 1,024 times the pixels of a 32px image! and if you are using quadratic attention, it costs even more than that), and the learning will be better because it learns global concepts before attending to the details.It also has any-time benefits for incremental sampling: you can decode the first few levels and use them, such as displaying them to the user at a low resolution, and then finish generating the final highest-resolution sample only for one. One might display dozens of low-res samples, and the user picks one—cheaper than diffusing all dozens at their most expensive resolution!And multi-scale representations are useful for documents which mix modalities: in a document format like HTML, we would like to insert images ‘in place’, like CM3, but we face a bit of a dilemma: the less lossy the inline image tokens are, the more it overloads our LLM and clutters the page. With a multi-scale representation, we can easily train on the full-resolution images separately, and then inline the thumbnail of each image—giving us the best of both worlds. We can do this with embeddings too, which might be useful for better conditioning or sampling (cf. “hybrid LLMs”).Taken to the extreme, one could incrementally sample arbitrary sets of pixels: if pixels are encoded as pairs of coordinates + pixels, then it can be trained similar to a MAE—simply append the coordinate of the next desired pixel and predict the pixel. This would be similar to crop conditioning. This offers the same advantages as a MAE of being able to train on very few pixels out of an entire image, which is where most of the computational expense of an autoregressive model comes from.-
multimodal planning: one benefit of a VAR multi-scale image representation (and for all other modalities where multi-scale makes sense) is that it is a good representation for summarization/retrieval/planning.
An agent LLM can ‘imagine’ the thumbnails to plan out and think in an inner-monologue, where the full-resolution tokenizations would be vastly excessive detail and bloat the context window, costing both compute and ability to intelligently use the context. A LLM agent like ‘Claude Plays Pokemon’ seems like it would benefit greatly from being able to put in little thumbnails copies of observed states, or as little representations of goals: “we are currently looking for a building that looks a bit like this” or “if I do X, does it look like this or that?” or “I tried Y last time, and the result was like thus and thus and thus and I’m not sure why.”
embeddings: Could the autoregressive approach provide ‘embeddings for free’ (like in iGPT)? If a fixed number of empty tokens are interspersed between the text/metadata prefix and the generated image tokens, and the Transformer attention patterns are restricted so the prefix can only attend to the in-between, and the image-tokens can only attend to the in-between, then that turns it into a bottleneck architecture like an autoencoder; and the bottleneck would presumably become an embedding.
-
-
Nenex: an integrated neural text editor/personal wiki idea (2023-09-13)
-
Hacker News Anecdote Search: the most valuable part of Hacker News comments to me are always the personal stories & anecdotes. Many are unique and will never be written about again. There is no source collating or curating them.
The problem with doing so is that the best HN comments are not always top-level comments, but buried in threads, and so they neither get read much nor do they get upvoted much. Nor is there any useful keyword for screening (what are you going to do, search for the word “I”?). You could try to do it by hand, but reading ‘all HN comments ever’ would take a lifetime or two. This renders them mostly invisible.
This is, however, precisely the sort of ‘I know it when I see it’ natural language task valuable requiring scale that contemporary DL NLP approaches now excel at. One could take an LLM or similar tech, curate a small set of HN comments on a few relevant dimensions (“funny”, “interesting”, “informative”, “historic”, “unique”, “war story”, “high quality” vs “low quality”) to finetune a classifier, and quickly bootstrap a large corpus of HN which can be sorted by dimensions and curated easily.
And it’d make a great “Show HN”.
“LLM Applications I Want To See”, Sarah Constantin
-
Hierarchical Embeddings for Text Search:
‘You Could Have Invented Transformers’ tutorial proposal (2022-04-06)
Generative Archiving
(2024-10-15) Can we develop a meta-prompt using prompt evolution/optimization frameworks for the purpose not of jailbreaking, but creativity?
Users of RLHFed/tuned LLMs badly affected by mode collapse, damaging their creativity for regular writing, have noticed that it seems you can reduce the damage by exhorting them in a prompt preface to try to be creative, to not be afraid of taking risks, and telling them that no one will judge or criticize them for trying. Past research on jailbreaking LLMs has generally found that the safety measures are “shallow” or “superficial”, and roughly equivalent to finetuning a few parameters or a long prompt, and so easily undone; this implies that such measures are not fundamentally changing the LLM (eg. it is not ‘erasing’ the LLM’s knowledge of naughty things), and simply a sort of specializing or biasing the LLM towards refusal or other ways of satisfying the criteria. Since these measures can be so easily undone even for the intended usecases they exist to eliminate, it should be easy to repair the damage on completely legitimate, harmless usecases.
So, it stands to reason that there ought to exist “creativity meta-prompts” which restore the base model creativity. How could we find one?
The many existing prompt tuning approaches should in theory be capable of this, but it’s unclear if they would work. They are usually designed for adversarial tasks, eliciting illicit outputs, which can be easily detected by heuristics or just any output which is not a refusal.
We take a meta-learning perspective: we define a “creative writing corpus” and finding a prompt which makes that corpus of samples more likely (total log-likelihood), yielding higher average performance over all individual tasks (writing excerpts) in the broader family of the meta-task (creative writing in general), thereby improving performance on future tasks drawn from the family (user requests for creative writing, of the meta-prompt + instructions).
To do so, we create a corpus of writing samples, drawing from many genres & forms, on many topics, possibly individually quite short, possibly freely-licensed, which all share the property of “not sounding like ChatGPT” (and not sounding like each other). Then we learn or evolve a single meta-prompt to increase the log-likelihood according to the LLM of all samples when prefixed with the meta-prompt.
If the samples are high quality, and the prompt optimization works, the most successful prompt should be one that restores the full distributional capabilities of the LLM, tilted towards creative writing, and undoing the mode-collapse.
(One possible failure mode is that this would also undo the benefits of the tuning in making the model “instructible”: one would no longer be able to describe the intended creative writing easily. If this is so, one can prefix “instructions” to the writing samples, describing them as prompts; these would not exist for human writing, but one could just use a LLM to generate reasonable “prompts” for each sample before attempting to create the meta-prompt.)
See the later “Learning to Reason for Long-Form Story Generation”, Gurung & Lapata2025.
(2026-01-17) One issue with learning a meta-prompt is that we will be hamstrung by the starting LLM’s capabilities, precisely because the prompt is short and RL is ‘superficial’. There may be a lot of missing implicit or tacit knowledge which is not easily written down in words the LLM already understands. We will only be able to find ‘easy’ tweaks, and we will struggle to evolve long-range strategies like careful planning, iterative revision, or extensive brainstorming.
This is partially because we cannot directly optimize by imitation learning of those strategies—they simply are not reflected in any training corpus. No creative author writes down all their thinking or dead ends.
How can we go beyond learning a single generic prompt on a static generic model? In particular, how do we avoid learning a relatively simple generic solution like a jailbreak prompt which avoids chatbot-style prose, but doesn’t actually get us to deeper thought & true creativity? We want to train the LLM to make good use of a latent scratchpad, similar to GPT-4-o1’s RLVR, but we lack a simple 0 versus 1 reward function—clearly, exact correctness can make sense in tasks like math or coding, but rewarding a model on whether it was able to exactly predict a creative output like a passage of Shakespeare is absurd.
I am going to extend Gurung & Lapata2025 and propose that likelihood maximization on tokens of creative writing can in fact train creative writing, just as it trains seemingly everything else; but that naive next-token prediction learns to produce uncreative writing because of a lack of compute in the prediction. You get outputs which are the LLM’s best attempt, but just look “creative-ish”. There is simply not enough compute in the forward pass of a single token to recreate the full computational process of ceativity and learn sample-efficiently, in the same way that a small LLM struggles to learn multiplication if it cannot cache or amortize computation over intermediate inner-monologue tokens or recurrent state.
We can treat the planning phase itself as something to RL: it is a long temporally-extended series of actions (tokens generated during the ‘planning phase’) whose optimal choices are unknown (because no one has written them down in the past) but which can be rewarded (model optimized) by the final performance on some objective criteria (log-likelihood of a creative writing).
So we could meta-train a LLM to do creative writing like this:
-
curate a large diverse corpus of unmemorized creative writing samples, like poems
These can be human-expert-written, but they could also be synthetic, eg. heavily filtered best-of-n samples from randomized prompts. (A fixed curated corpus anchors optimization in a diverse immutable set of examples and so discourages mode-collapse or entropy collapse, but if using synthetic or self-generated data, this danger returns. However, across a sufficiently large and diverse corpus, shortcuts like “guessing the author” becomes indistinguishable from “learning the distribution of high-quality writing”.)
have the original LLM write a (short, length-limited) analysis & summary of each datapoint, describing its genre, style, goal, meaning, historical context etc
-
begin evolution strategies DRL:
mutate k copies of the original LLM
-
for each of the k LLMs:
prompt them with a datapoint’s summary and the instruction to “plan out how you would write this story between
<thinking>/</thinking>tags”roll out the LLM until the
</thinking>end-of-planning tag (possily annealing greater length)append the actual datapoint to this rollout and score the average log-likelihood of the datapoint (not the summary or planning phase), as the mutant’s total episode-level reward
keep the LLM mutant(s) with the highest reward
repeat
Evolution strategies is computationally expensive, especially at LLM-scale. Why that instead of a policy gradient like PPO? I choose Evolution RL because aside from being highly parallelizable, evolutionary RL algorithms tend to be excellent at learning temporally extended strategies with difficult-to-see whole-episode benefits in sparse-reward problems, while avoiding deceptive local optima or spurious gradients, and are noted for their creativity. Further, it lets us freely complicate the reward function and environment. We try to evolve a LLM which will meta-learn how to take a summary and on the fly, ‘think about it’: this DRL loop gradually trains LLMs directly to make good use of a penalty-free planning phase in order to most usefully predict (= generate) the final draft of a piece of creative writing. The more it taps into things like ‘try to think up relevant allusions’ or ‘what vocabulary would such an author use’, or writes out ‘drafts’, the better it can predict the true datapoint.
And this can be easily extended with the many ES improvements like sparse/low-rank mutation (eg. of a LoRA or soft prompt), antithetic sampling or to permit Quiet-STaR-style planning during prediction of the datapoint (just mask out thinking tokens from the log-likelihood), penalizing similarity to chatbot-like completions (ie. reduce reward based on log-likelihood of mode-collapsed datapoints), novelty search, adversarial tests where memorization won’t help (eg. procedural constraints, forced weird topics, rhyme schemes with new content), hard negative mining / dropping datapoints where the likelihood is too high (implies memorization or leakage), etc.
The final LLM can be distilled into another LLM or the original LLM (to fix damage/drift), by generating a final set of [summary, planning phase, datapoint] sequences and simply doing sequence-level distillation.
One might wonder to what extent this procedure is hamstrung by the quality of the original summaries: if they miss something important, or are wrong, then the LLM cannot meaningfully improve past a certain point. (For example, a LLM might know somewhere deep down about the religious allusions in a datapoint, but be too risk-averse talking about religion to include it in normally prompted summaries, and so how could the planning process ever help predict the datapoint? The “religion tokens” would just come as a perennial surprise.)
But we can just evolve summarization as well, by almost the same procedure, since “good summaries which are useful eventually to predicting a passage” is just the reverse process with similar DRL properties. In this case, we mutate the LLM doing summarization (presumably the same LLM we have been meta-training), and we reward the mutants based on how well the different mutated summaries increase its own log-likelihood of the final passage.
We alternate these optimizations, EM-like: planning phases are meta-trained given frozen summaries, then summaries are meta-trained given frozen LLMs executing planning phases, and so on. As the LLM develops ever greater creativity, its summaries become more and more insightful, which enable more powerful planning phases, which then enable more acute judgment of what summaries captured the key things. (This alternating optimization bears a pleasing resemblance to how human creative writers learn: interleaving criticism with generation, and asking, “how did they write that?”)
These summaries, being an autoencoder bottleneck in a mutually-evolving system, need not remain human-readable. This could be a feature or a bug—human creativity is often not explainable in words, and a ‘token soup’ might allow for generating random ‘interesting’ samples. If it’s a bug, it can be fought by the usual tricks like length limits, paraphrasing, restricting token vocabulary, requiring a strict format, swapping in other LLMs to break emergent conventions etc.
Wanted: better SVG generative models! Vector image generation models lag behind raster image models by a lot, as of January 2024. If one wants a good vector image, it works better to generate a raster image and then use a vectorizer to convert it.
Why are vector images harder? It seems to be hard for two reasons: vectors, to be useful, tend to be highly symbolic and modularized and programming-language-like.
It is trivial to convert a raster image to a degenerate vector image which has a vector per pixel, or to do something similar like differentiably optimize thousands of vector curves into an approximation of the pixel image, but this would generally defeat the point: it wouldn’t upscale/downscale well, it wouldn’t have meaningful semantic chunks like backgrounds which could be deleted or which could be edited by a human. (Whereas a pixel image can be made out of blobs and textures smashed together until they look good, with no expectation that the image would be made out of discrete modular parts.) Quite aside from the basic challenge of writing out thousands of SVG text tokens without any syntax errors or other issues, writing out a good vector image seems more challenging, and on a higher semantic level, than a raster equivalent.
Secondly, and related to the previous difficulty, complicated high-quality vector images are hard to make and scarce compared to raster images. Anyone can pick up their smartphone and shoot a high-quality photograph of something in front of them, like their cat; but to produce a vector illustration of said cat is quite another thing. Even a simple UI icon for a web page or an app is a fairly specialized skill. So, it’s no surprise that while it’s easy to scrape billions upon billions of images from the Internet, vector datasets remain in the low millions, generally, and have severe quality & diversity issues. (Almost all of them will be icons, for example, while the ones which represent large complex scenes or illustrations may be poorly auto-converted from raster images and liabilities for training.)
And if there is a dearth of SVG files, then there are even fewer SVGs with adequate text captions describing them. So we are in a bad place for any ‘text2vector’ generative model. We have so many raster images, many raster images’ text captions, some SVGs, and just a few SVGs’ text captions. And we want to go from the least common to the second-least common type, where that type is also an intrinsically difficult type of data.
It is no wonder that vector generative models tend to be highly limited, low-quality, or aim at other tasks like converting raster to SVG.
My suggestion for how to handle these problems is to generalize & scale it: create a single generative model which handles all those modalities and translates between them, treating it as an autoregressive sequence modeling problem, like DALL·E 1 & CogView.
The original DALL·E was a GPT-3 which trained on sequences of text tokens, concatenated with image ‘tokens’. So it learned to take a text description and predict the image tokens of the corresponding image. CogView noted that there was no reason that you could not simply reverse this: instead of [TEXT, IMAGE], [IMAGE, TEXT], which is an image captioner. And since this is the same domain, one can use the same model for both and share the knowledge. (Just train it on both ways of formatting the data.) And one could keep on adding modalities to this. One could add audio of people describing images, and train on [AUDIO, TEXT, IMAGE] + [TEXT, AUDIO, IMAGE] + etc. Now you could record someone talking about what they see, predict the text caption, and then predict the image.
This is particularly useful because some modalities can be easier than others: we may have much more of one kind of data, or be able to construct pairs of data. We can apply various kinds of roundtrip techniques like backtranslation or data augmentation, or exploit synthetic data to reverse a direction. (For example, we could feed a dataset of text into a voice synthesizer to get audio descriptions to train on.)
In the case of vector generation, we would train on all permutations of [SVG, raster image, text caption]. The benefit here is that it covers all of our desired use-cases like SVG → image or text → SVG, and enables bootstraps and synthetic data. For example, if we have a random unlabeled SVG, we can still train on it directly, and we can also fill it out: given an SVG, we can create its raster image easily, and we can then use an image captioner to caption the raster image. Now we have all the pairs.
We can improve the generator by roundtrips: SVG → image → SVG should yield near-identical SVGs, and vice-versa. We can especially exploit synthetic data: we can superimpose SVG images on top of each other in specific relationships, and learn to generate them unconditionally (which would encourage the generator to learn to cleanly write separate objects, even for extremely complicated or cluttered scenes), render them into images and train to generate the original SVG, or generate a text description and SVG (and rendered image) all together. For example, one could imagine a little Euclidean domain-specific language, perhaps like the infamous SHRDLU ‘blocks world’, which generates sets of blocks like “a triangle on top of a cylinder to the left of a sphere”; one can render that scene as an SVG, and this would help with the sort of relational reasoning that existing generative models often struggle with when you ask for ‘A to the left of B’. This can be arbitrarily augmented: add an SVG object of a growling lion, for example, and now you can have “a lion behind 3 blocks”. We can use all sorts of transformations which ought to commute or be identity functions—eg. corrupt the SVGs or images, and roundtrip through an image. (‘Corruption’ here can include lossy transformations like upscaling & downscaling: if I resize a 1024px PNG to 256px, they should yield nearly-identical SVGs, and if I grayscale it, it should yield something perceptually & textually similar to the SVG of the color image which has been grayscaled in SVG-space.)
Or indeed, why not train on more than one ‘SVG’ modality? We can define arbitrarily many ‘kinds’ of ‘SVG’: ‘SVG’ as generated by specific tools like potrace, or after going through an optimizer/minifier tool, or generated by previous versions of the model. (Simply prefix all the metadata to the tokens.) So one could train the model to generate [SVG, minified-SVG], or [syntactically-wrong SVG, SVG]. Broken or bad SVGs can be generated automatically by adding noise to SVGs, but also collected from the model’s errors, both during training, and if a user submits an example of a text prompt which yielded a bad SVG and the intended good SVG, one can train [text, good-SVG] but also [text, bad-SVG, good-SVG].
Obviously, the rendered images provide useful training signal to the model trying to generate SVGs by informing it how the SVG should look, but because we can encode more metadata, we could go beyond simply presenting the pairs or triplets for autoregressive model. We could, for example, include human ratings of image quality, as is now standard in preference-learning. But we can go even further, and use a pre-existing image model like CLIP to add in metadata about how bad an image looks, how far away from the intended image a generated SVG is, by running CLIP on the rendered image of that SVG versus CLIP on the ground-truth image, and then encoding that into the metadata to condition on. (“Here is an SVG, which yields an image which doesn’t look anything like it’s supposed to, it’s off by 0.35 in embedding space. On the other hand, here’s an SVG which looks nearly-identical to what it’s supposed to look like, a distance of 0.00.”) This might be particularly useful for tricky SVG tasks: if we generate a hundred possible SVGs, and they all fail to yield the target image, that’s not much of a training signal; but if they all turn into new training data, and they all include an objective absolute measure of how far away from the target image they were, that is rich feedback which will improve future attempts; and this could be done repeatedly. (This would potentially enable the model to slowly bootstrap itself up to a specific image, by trying out many possible SVGs, training on the results after being run through the SVG renderer, and trying slightly better SVGs the next time, until it ‘predicts’ an SVG distance of 0.00 and is correct.)
LLMs currently do not write competitive novels as of June 2024 (that we know of). RLHF & safety-tuning aside, even the ‘base’ LLMs struggle at writing high-quality plot-heavy novels of the sort which can span hundreds of thousands of words.
This struggle is despite the fact that the main technical barrier—too narrow context windows which simply can’t fit much text—has largely been lifted at this point, with LLMs like Claude-3 Opus, ChatGPT-4o, and Gemini Ultra capable of processing context windows with hundreds of thousands or even millions of tokens; so in theory, an entire novel, or even multiple drafts of a novel, can now be held simultaneously in an LLM context window.
The main issue now seems to be general coherence of planning: writing a good novel in a single sitting without any kind of rumination or thinking is not something any human can do—no, not even the “pantser” serial fiction writers (who will draw on a store of ideas & fragments, and whose minds will keep pondering their in-progress work in between each bout of writing).
And that lack of planning ability must be in part a lack of training data about how to write a novel. Because while there are millions of novels to train on, yes, almost every single one of them is a finished, edited novel. In reality, novels are of course typically written with a much more extended process—months or years of planning, pondering, background knowledge, outlining, writing, revising, editing, rearranging, and so on. Hardly any novelist simply sits down and types out an entire novel which needs no editing. (The ones who do are typically writing pulp garbage or highly unusual kinds of writing, like graphomanic stream-of-consciousness.) A writer like James Joyce did not write Ulysses or Finnegans Wake in a single sitting, but instead painstakingly revised it for decades.
But this process as a whole is invisible, and not represented in training data to a meaningful extent; there is no ‘Github of novels’, nor can authors write down the whole process to begin with. No one writes or publishes a detailed outline of their novel, from the top-level down to each page, before they write the novel itself, and they don’t publish notebooks full of random ideas or discarded concepts either.
So it is unsurprising that LLMs do not do the process well. (Adding in adaptive computation techniques like Quiet-STaR would presumably help, but are still no panacea.)
But as with the hierarchical RL approach, every novel represents an answer to the question of writing a novel, and we can reverse each finished high-quality novel into a hypothetical process which yields that novel as its result, and we can modify that process to stress the LLM in various ways. (“Every solved problem is also the solution to a harder problem.”)
LLMs can recursively summarize novels, even as far back as GPT-3; a hierarchical summary of a novel, going paragraph by paragraph, page by page, chapter by chapter, section by section, and summarized as a whole, is a ‘planning process’ for writing a novel as well. Simply format the data in reverse order: the shortest highest-level, followed by sections, then chapters, then pages etc. (eg. my Claude Rubik’s Cube essay.) And they can then be prompted to revise sections repeatedly, to improve each one. (eg. my Perished Paradise examples.)
LLMs can be prompted to generate works in this way, like DeepMind’s Dramatron, which uses hierarchical prompting to run summarization in reverse: recursive expansion.
This logic can be extended. Recursive expansion doesn’t spend any time “brainstorming” or “thinking”; there are no “writer’s notebooks” for LLMs. But we can synthesize such things too.
We can ask a LLM to analyze a novel, and pull out ideas, like “tropes” from TvTropes, and make a long list of interesting places, people, ideas, tropes, plot events, and so on, for each novel. (This could be done as part of the recursive summarization, if each summarization includes a list of interesting points, and that post-processed out for a clean summary.)
Then we can synthesize a brainstorming phase by, for each novel, creating a “list of ideas” which is all the original actual ideas for that novel, randomly mixed with a bunch of items from other similar novels. Then appended to the long list of real+fake ideas, is the final clean actual list for that novel.
This helps teach the LLM ideas & concreteness (“show, don’t tell”) and avoiding the vagueness of typical LLM sampling, but also creates a brainstorming phase when it is run forward: the LLM comes up with a big set of possible ideas, and then picks out a coherent subset, and uses that in the following recursive expansion.
An open question: if the revisions are not good enough, how can we synthesize low-level editing of passages? We still don’t have the revision process of the actual low-level text, even if we have synthesized the planning process. We only see the final text, purged of all noise & errors in any original draft. And ‘noise’ in text has no simple statistical equivalent to Gaussian noise in diffusion image generation models which we can easily synthesize, so how do we do a backtranslation/reversing process if we can’t go either direction? (Discrete diffusion models capable of text generation exist, but remain highly experimental and underperform autoregressive LLMs.)
It is possible that a LLM could be prompted to “make this passage worse-written”, given that LLMs are capable of writing in all sorts of styles, at all levels from college to “ELI5”; so one could repeatedly degrade a passage, and then simply reverse the sequencing with “edit this passage”, and that forms the training corpus.
Another possibility is to exploit the hierarchical expansion: generate “worse” versions of a passage by simply conditioning on less—drop sentences from the summary being expanded, particularly omit ideas from the list of ideas, and then ‘edit’ it into a version with more specified details. So the ‘worst’ version gets the shortest summary, with all the ideas deleted; then a second not-so-bad version is generated by prompting it with a larger summary plus some ideas, and an ‘edit’ prompt; then a third better version by adding some more back in and editing, and so on; then the final dataset is simply the full prompt → bad version → not-so-bad version → … → final best version. This maintains semantic similarity, and creates a sort of Brownian bridge steering towards the intended final version (diffusion at a more conceptual level).
A third approach would be to use self-distillation/rejection sampling: generate n completions, use the LLM to pick the best and worst completions, and then create a synthetic “edit” datapoint by worst → best.
In a base LLM model like GPT-3 or GPT-4-base, which do not have the default ‘assistant persona’ of RLHF/instruction-tuning/RLAIF like ChatGPT or Claude, one can usefully do writing by setting up a ‘simulation’ scenario with a persona or character. The user can instruct the persona to generate a particular kind of completion, to steer the results in desired directions.
The straightforward approach yields only one completion from the persona, at which point the dialogue presumably passes back to the user for a response; rather than use some sort of external ranking approach on a bunch of sessions generated in parallel, one can instead ask (within the narrative) for multiple completions, as alternate versions of the story, and then ask the persona which one is best. This yields something like a best-of-n likelihood ranking result, as seen by the LLM. (It’s unclear whether it’s better, worse, or just different from best-of-n.)
This has a couple downsides:
it is limited by the context window (often quite small for base models): if all n results do not fit in the context window, the persona can’t see and pick from them
-
it provides no information about the unchosen alternatives, so it is weak for search purposes; it is hard to infer all the ranks
If we wanted to score them all or just rank them, we would have to try something more complicated, like treat it as a sorting problem and use a lot of pairwise choice completions to infer a noisy sorting—requiring 𝒪(n · log(n)) completions in the best-case scenario of little LLM noise, but possibly a lot more.
-
the narrative and instructions risk ‘framing effects’, distorting the completions:
for example, if we add an instruction like “Choose the continuation that [x].”, this unavoidably carries a whiff of a Choose Your Own Adventure text game or an academic exam, which may be highly undesirable framings.
-
there is an inherent but here difficult-to-control tradeoff between the competing objectives of the quality of the completion and the compliance with any instructions
There is no straightforward way to change the tradeoff, because any sort of natural language instruction which tries to say “X is very important!” or “if possible, I’d like X”, might be interpreted in wildly different ways from session to session or model to model or not form a smooth progression etc
-
potential harmful interference between completions:
the more ‘completions’ in the context window, the more the base model may incur few-shot conditioning problems, and lose diversity or begin repeating or fall into repetition traps (Mode collapse is particularly risky for anything creative or novel.)
We can reduce this by instead sampling in parallel sessions, and then constructing a session for a final persona ranking, but this adds complexity.
(Done right, ‘interference’ could be useful, if that can guide the model to generate maximally different completions, but it’s unclear how well models can do that or how to naturally prompt them. Is it enough to instruct the persona to “tell a different story each time”? Do we need scaffolding like “list 10 ideas for stories”?)
We can’t use best-of-n likelihood ranking as-is, because it doesn’t directly take into account the instructions, nor can we just compute windows over the total completion.
Because base models often come with likelihood of tokens, I propose an alternative: we can combine the original Meena best-of likelihood trick with the inference-as-compression paradigm where compression estimates K-L distance.
The original 2020 OA API offered a (now long-deprecated) “classification” function, which let one input a list of categories (text) and a target text, and returned the similarities. This did not use the usual approach of embedding, because decoder-only LLMs like GPT-3 do not inherently generate useful embeddings and OA had no embedding-generating GPT-3 model at that point. So what did it do? What it did was see which combinations of text were more predictable. If text B can be more easily predicted than text C by GPT-3 given text A as its prompt, then text A can be ‘classified’ as a ‘B’, whatever that is. More easily predicted is an absolute concrete likelihood number, so one can do this for an arbitrary number of texts, and the texts are fixed and it works with any text. (This is an old compression gimmick: one can, for example, use gzip to concatenate pairs of text files and do things like infer phylogenetic trees from compression ratios.)
We can do this for completions too. We can go back to generating n separate completions, and then append the instructions to the completion and see which ones “compress” better. The more closely that a completion follows the instruction, the more the instruction will be compressed, because if you see something which is an example of an instruction, the better it is an example, the more you can guess what the instruction was. This gives us two likelihoods for each completion: the likelihood for the completion, and the likelihood for the instructions given the completion.
This avoids several of the problems of the in-band ranking:
we can now rank completions which fill (almost) the entire context window, since we are looking at one at a time (plus the instructions)
no framing effects: because the instructions come after the completions are fixed, the completions cannot be affected (and since the instructions are themselves fixed, there’s no framing issues there either)
-
we can explicitly tradeoff the two likelihoods:
If we care solely about quality, we look at just the former and zero the latter; if we care solely about instructions, vice-versa; if we want anything in between, like ‘high quality with some weak nudging from the instructions’, we simply take a weighted sum. (And we can generate a Pareto frontier of samples, if we have a lot of them.)
-
we have absolute scores, which give us a complete ranking—difficult with in-band ranking.
-
we can rank an arbitrary number of completions: given the scores, we don’t need any additional completions at all.
So instead of paying 𝒪(n · log(n)) comparison-completions to rank n completions, we simply pay for n completions.
we can use these scores to guide a tree search (like a novelty search), and provide feedback to the user about confidence or quality. (‘Search’ might be as simple as the user rejecting the top few for various flaws and then keeping the highest remaining.)
-
independent samples can’t affect or contaminate each other, and can be sampled at different temperatures or high temperatures
The downsides are that this is substantially more complex, and also comes with a substantial overhead in completion size which may not be amortized over enough completions.
So these approaches could be combined. I would guess that the persona choice might be ‘better’ than the likelihood ranking, so one could create a hybrid approach: use the likelihood ranking to construct the full ranking in an efficient scalable way, and then take the top k samples (whatever fits in the context window) and then ask the persona to choose. Given a reasonable bivariate correlation, the top k by likelihood will be highly likely to contain the best by persona choice as well.
Our historians, the most perspicacious on the planet, have invented a method for correcting chance; it is well known that the outcomes of this method are (in general) trustworthy—although, of course, they are never divulged without a measure of deception. Besides, there is nothing so tainted with fiction as the history of the Company… A paleographed document, unearthed at a certain temple, may come from yesterday’s drawing or from a drawing that took place centuries ago. No book is published without some discrepancy between each of the edition’s copies. Scribes take a secret oath to omit, interpolate, alter. Indirect falsehood is also practiced.
Error-Ensuring Codes (2024-09-01): a proposal for a generation method to reduce the harms of synthetic media by deliberating introducing minor errors which are locally consistent but globally inconsistent, and relatively cheap to detect but computationally intractable to remove.
This would be a useful form of watermarking for AI outputs not intended to be strictly factual, like most image/video generation uses.
A common proposal to reduce the harms of generate models creating floods of meaningless but superficially true-seeming media is to do “watermarking”: add some sort of subtle statistical signal to the unimportant details of the generated media, like the slightest shades of color or unimportant word choices. This watermarking is designed to be easy to detect (at least by authorized parties) and unforgeable (so people can’t discredit real media).
However, watermarking has the issue that for pragmatic, utilitarian purposes, watermarking is neither necessary nor sufficient: completely legitimate media can be watermarked, while the useful content can usually be extracted & rewritten as a kind of ‘analogue hole’. (This is also true of approaches to try to detect neural signatures like that of top-k sampling.)
This sort of watermarking is useful for catching cheating students, or possibly lazy employees violating corporate rules on use of generative models, but it is not too useful.
A bigger harm of synthetic media is that it creates realistic, self-consistent, fictional worlds which cannot be distinguished from real-world data—like Borges imagines (“Tlön, Uqbar, Orbis Tertius”) a secret society of geniuses conspiring to create over the centuries, which poisons the real world, and slowly takes over. It is poisonous because its verisimilitude ‘pollutes’ commons by leaking out of its intended uses and being mistaken for true real media even though random synthetic media cannot be any more true than its sources or otherwise add value.
This has been a long-standing problem with media, especially special effects, where what was never intended as anything but entertainment (often satirical), gets pulled out of context and presented as a genuine; but gets far worse with generative models & social media, where anyone can fire up a free image generation service and immediately mislead millions (particularly as context-free copies outrace any kind of factchecking). Such images are rapidly approaching indistinguishability even for those long familiar with AI artifacts or the highly-idiosyncratic styles of mode-collapsed image models like DALL·E 3.
Worse, we are increasingly seeing multimodal synthetic media: like websites where all text, photos, illustrations, JS/CSS/HTML, associated Github repos, and even videos are synthetic, and all mutually consistent and self-reinforcing—but fictional (because in the service of a spammer, scammer, hacker, nation-state, art project, marketer, etc.). Projects like Websim are just an early taste of this.
This raises a serious problem for future factchecking, including AI-powered factchecking. Most frauds are ‘shallow’, in the sense that if you do any level of verification and trace claims to their roots, the imposture quickly becomes evident: the company doesn’t exist in the business databases, or the address is of some other business, or it cites a book which doesn’t exist or where the page does not support the claim. When a copyright fraud extortionist sets up a fake law firm for the shakedown attempts, they may register a domain name, generate StyleGAN headshots, and use a LLM to send the emails or write the law firm website’s text, but they generally do not go any further; when an industrious Chinese Wikipedia user wrote hundreds of fake Russian history articles, she simply made up the citations to real books, and so when people finally started checking her citations, the fakery was exposed immediately. Historically, if one traced a claim which turned out to be bunk, like “Egyptian tomb honey”, it is usually not that hard to debunk it or find the (very) small grain of truth behind a famous ‘fact’; checking citations 1 level deep will work a depressingly large fraction of the time due to miscitations & everyone being too lazy to check even that much, but if not, usually going 2–3 deep is adequate. (The challenge with these ‘leprechauns’ usually tends to be getting the obscure offline sources at each level.)
This is because even a determined fabricator usually cannot fake entire journals or articles or authors or websites recursively, so they cannot achieve “epistemic closure”, and at the ‘edge’ of their citation graph, they are forced to either make up the claim without citation/sourcing or risk a miscitation by misattributing it to an innocent real source. (A case like John Drewe, where he focused his art forgery efforts on all the support documentation, while the art itself was trivially debunked, is an exception which proves the rule.) So simply going depth-first will rapidly expose it.
Unfortunately, with generative models, it will become entirely possible to close the loops: claims will dead end in what seem to be real papers written by real people with real photos and real data, who did make the claim as claimed by the most proximate source. And there will be plenty of incentive to do so: TV shows, novelists, and video games exult in worldbuilding—the more detailed and realistic, the better.
At that point, the fact-checker is in serious trouble: it is no longer enough to simply trace a few citations from the comfort of one’s computer. Nor is it enough to simply verify a watermark indicating that some text passed through OpenAI’s servers at some point—if the document is from pre-watermarking, then sure, that disproves the truthfulness of the document, but otherwise, it means little. One may have to start doing real-world things to establish that authors didn’t exist or the events did not happen, and proving such negatives is both a lot of work and itself highly unreliable.
Then the fake world will be pulled up into generative models by later scrapes, and start influencing all future generations, in a much more worrisome fashion than simply low-quality random AI samples.
What do we need? While we cannot hope to stop fake synthetic media abuse in general, we can hope to reduce the abuse and exploitability of synthetic media create by responsible actors. We need some sort of ‘watermarking’ which does not indicate who processed some media, but that it is unreliable and should not be believed as factual or truthful (although it is fine to enjoy it as entertainment or use it for many other purposes).
We’d like a honeypot like a trap street or the 555 telephone area code, which tells you that it’s not a real telephone number and doesn’t really go to where you see it despite being captured on camera, and whose inclusion tells you that you are dealing with fiction.
But changing the 555 area code in every depicted telephone number to a real area code would be pretty easy, so it’s not enough to adopt a convention like that; the ‘falsehood watermarking’ should ideally not be fixable by ordinary anti-watermarking tricks like local rewording.
Since we don’t need to worry about preserving the truth or semantics of the synthetic media (it has none to begin with), we are not restricted to the usual subtle modifications like the most invisible pixels or subtle wording choices.
What we want is sort of the opposite of an “error-correcting code”: an error-correcting code ensures that the original coherent, meaningful, true data can be recovered exactly even if a lot of it has been corrupted or deleted, when we take a global view; here we want the opposite, an “error-ensuring code”—something which ensures that coherent, meaningful, true data cannot be recovered even if a lot of it is deleted, that there is always some tell-tale of being fake, some artifact or error or contradiction. Then downstream users can do basic sanity checks and detect the problems, and know to investigate more deeply, or throw out the data. (They don’t care who exactly generated it, only that it was generated to be fictional and should not be trusted or trained on naively.)
What could an error-ensuring code be?
Here’s one idea: by analogy to probabilistically checkable proofs & the PCP theorem (cf. chaffing and winnowing, chaff bugs), the generation process can insert errors which are locally consistent but globally inconsistent. This is because it is easier to create or detect some errors than it is to fix all errors.
It could, for example, use prompts or ‘knowledge editing’ procedures (like ROME or activation steering or SAEs) to cryptographically randomly change what the model believes periodically throughout the generation process, at minimal computational cost.
For example, it might believe for the first half that ‘Paris is the capital of France’, and then believe for the second half that ‘Marseilles is the capital of France’, or that George Washington was elected in 1788238ya, then 1789237ya, then 1790236ya. And it might use straight ASCII quotes in one part, but then Unicode curly quotes in another. (And similarly for other modalities: video or images might use slightly different backgrounds, or different times of day, each equally valid in isolation, but jointly impossible.)
While any individual bit of knowledge might not affect a passage, a large enough set will create subtly inconsistent ‘worlds’ that each passage is set in, ensuring the pervasive presence of errors when the passages are compared. (These ‘inconsistencies’ may not even be perceptible by a human, given LLMs’s situated awareness & truesight: a LLM may be able to detect that ‘different’ LLMs generated these shifting passages due to stylometric-level distortions but not know which or localize the exact telltales which cumulatively constitute the textual signatures.)
In an image or video generation context, an image can contain physically-impossible inconsistent geometry, which is not noticed by a viewer simply passively consuming it and which may in fact be desirable for its “look”, like contradictory shadows; simply spot-checking a few random pairs of shadows or objects may immediately reveal contradiction and thus fictional. (But since the geometry is so important to the overall scene, it may not be possible to easily repair an image without making esthetically or semantically-dangerous edits like erasing objects or extending shadows, which may damage its composition or change its meaning.)
A downstream user can then detect errors using LLMs by simply randomly sampling pairs of passages to look for any contradictions, in addition to straightforward errors. The more consistency checks the corpus passes, the more reliable it is. (And to the extent that this mistakenly flags genuine media which happens to be rife with errors & inconsistencies in an all-natural fashion, then the damage should be minimal as such media should be avoided.)
Meanwhile, an attacker trying to repair a large corpus to remove the error-ensuring code faces the problem that while they too can randomly sample and attempt to fix those contradictions, there will be countless more contradictions to find, and worse, the fixes themselves will introduce more contradictions.
Trying to repair all these contradictions simultaneously corpus-wide may require iteratively solving an extremely difficult global constraint satisfaction problem, as each fix causes new problems which need to be detected & fixed. An analogy might be the Ising model (visualization): finding the minimum-energy, which makes all the magnets line up consistently, by flipping magnets one by one, is intractable because any flip to make a magnet consistent with one neighbor may well just make it consistent with others—if one can flip all magnets simultaneously, it can still take many, many iterations (with each magnet flipping repeatedly) for all the ‘waves’ to propagate and the magnets as a whole to finally reach a consistent configuration. (I suspect that, like Ising models or satisfiability problems, there is probably a phase transition in the number of claims, where a self-contradictory corpus transitions from being computationally easy to fix to being computationally intractable.)
At that point, it will hopefully be easier for an attacker to give up and go generate a brandnew corpus, instead of freeriding on existing ones.
This would mean that self-consistent text at least serves as a costly signal or proof-of-work: either someone wrote it by hand as a human, they invested a large amount of compute to clean up an error-ensured corpus, or they were forced to use a more expensive / worse quality AI model without safeguards.