One might wonder: is an AUNN truly the very simplest possible NN architecture?
Maybe not. We still have the index input making it more complex.
Do we truly need indexes? It is noteworthy that strictly speaking, Transformers do not need ‘indexes’, in the form of positional embeddings, at all. Causal language models still work well without them. (Not as well, of course, but a lot better than you’d expect.) The causal language model analysis shows that the causal language model implicitly learns to encode the position of tokens based on distance from the missing token.
This is initially surprising, but on closer examination makes sense. Often, positions can be reconstructed and a sensible prediction made even if the input is randomly shuffled or is in a set-like representation.
It is clearly possible in principle, if you know what you are doing, to receive corrections for the previous token, and then make a prediction for the next token: humans can do that easily. If I show you the sorted input ["The", "brown", "fox", "jumps", "lazy", "over", "quick", "the"], with a little thought, you can easily predict that the next token is “dog”. And if you predict wrong, you will be able to make a better guess about the next token too. You can change your next guess based on the feedback for your previous guess. (“What is the next token?” “I think the next token is ‘over’.” “Wrong: it was actually ‘the’. Now what is the next token?” “Oh. Uh… ‘lazy’.” “Right. Now what is the next token?” “‘Dog’.” etc)
Nor does this seem implausible for an artificial neural network to be capable of doing. All this requires is that gradient descent be able to copy a datapoint into a model in a single update (which we know is possible from NNs being able to memorize data they have seen only once), and for the model to be able to think about a datapoint it has memorized (which they must be able to do in order to do anything which is not explicitly mentioned in an input).
So, do AUNNs need an index input? How could an AUNN work if we just drop the index inputs entirely?
Well, we’ve already seen how this can work by invoking neuroplasticity & dynamic evaluation! We can move further away from memorization and double down pure, continuous online learning for everything, even without any explicit memory mechanisms. We simply need the index to be implicit and inferred during training. (Arguably, this is what a dynamically-evaluated ICL-heavy model like TTT RNNs or RWKV-7 is doing already, when the ‘outer’ network calls the ‘inner’ network, and perhaps is like an extreme Belief State Transformer.)
We remove the index (let’s call this AUNN variant an Input-Free NN or IFNN), and we change the training scheme to simply proceeding through the data, which is in a coherent (usually temporal) order, and we take a single gradient descent step for each token to train the IFNN to predict that token. Hence, an IFFN is an arbitrary feedforward blackbox, with no input, which outputs a single binary variable, and our AUNN reduces to just this:
The full IFNN architecture (schematic diagram).
Of course, the token changes each time. So how does this work?
Well, repeated over enough data, with enough model capacity, what we are teaching the IFNN to do is to engage in online learning, and to find a meta-learned model which rapidly updates first-order REPTILE-style “where in the data it is” and based on that, the gradient descent updates it to predict based on the training target, the next token after that. As before, we expect some of the weights to automatically specialize over the course of training to become ‘fragile’ and the equivalent of self-attention or hidden states, encoding the sufficient statistics for a task, and those weights are what is mostly changed in the gradient descent step. (Language models already are non-myopic: they must implicitly plan and reason about future unseen tokens in order to more accurately predict the current token, and this has been confirmed by mechanistic analysis; we just incentivize that more.)
This won’t be easy, because our standard first-order SGD is aimed directly at the wrong thing: predicting the (now stale) current token. It can be quite hard to train even normal neural networks, and it’s taken many decades to discover current deep learning recipes and that even the safest-seeming choices like sigmoid activations are dead-ends. So I do not claim that IFNNs are easy to train, just theoretically possible. Our NN may have to be trained for a long time in order to be able to update in just the right way. This is something that might be better off done using a second-order gradient descent algorithm (like the original MAML, rather than REPTILE), or perhaps using an evolutionary algorithm (which can ignore deceptive gradients entirely and travel directly towards optimizing the ‘next-next token loss’, as it were). One could experiment by training IFNNs which predict multiple tokens in parallel, rather than just one, which should be more Transformer-like and stabilize training; and gradually anneal to a single output.
OK, so that, weirdly enough, trains a generative causal language model.
And once you have that, you can then engage in the usual prompt programming, encoding of tasks and delimiters, multi-modality tokens to interleave images and text, reinforcement learning by imitation learning, and so on and so forth—as long as you always ‘sample’ from the IFNN by generating and then updating the model, to increment the implicit index. We are not constrained to simply sampling forward either; like an RNN, we can run forwards or backwards or in arbitrary ‘directions’ across our ‘data’. The ‘temporal index’ is about the gradient descent, and we do have to update the model serially in time, but there is no restriction on the data, and we could go backwards through an list and then forwards as we please, as long as the IFNN can predict the changes in direction and avoid a catastrophically destructive update.
So an IFNN can do whatever you want.
What are the advantages of an IFNN?
-
it is pathologically simple
It is unclear to me how a NN architecture could be simpler even in theory. (How could one use a NN which has neither inputs nor outputs?)
-
may have better inductive biases towards time-varying dynamic systems, including ones where the system itself changes, because of its own path-dependence and not forcing any a priori distinction between ‘state’ and ‘weights’
free anomaly detection via gradient spikes?
-
much more flexible updating patterns: things like bidirectional RNNs or recursive NNs have fallen out of favor due to not being GPU-friendly and being trickier to implement; however, for an IFNN, every update pattern is the same.
This might be the best part of an IFNN. An AUNN requires careful thought about the layout of indices, to avoid clobbering already-used indices, and wants you to go in specific access patterns. But an IFNN may be happy to pivot instantaneously in any direction and forget about stale data as quickly as possible. One may be able to casually dump in data and change tasks on the fly in ways hard to foresee.
more parameter-efficient as no longer incentivized to memorize as heavily as an AUNN
-
? ? ?
-
we could wonder if the online-learning-by-neuroplasticity paradigm could have some sort of biological interpretation or have unique scaling behavior, which contrasts with the AUNN “memorize the world” strategy which is highly biologically implausible.
The idea of learning point by point purely by neuroplasticity is certainly reminiscent of how organisms must learn; they cannot stop and train their brains for a few months by repeatedly doing a big minibatch over data randomly sampled throughout history.
Perhaps an IFNN would become extremely sample-efficient and adept at meta-learning in general because it is trained solely online; or perhaps an IFNN could generalize unusually well and be adversarially robust, as the training implicitly heavily regularizes it?
-
What are the disadvantages of an IFNN?
-
training is inherently serial as an IFNN is fundamentally path-dependent.
I cannot immediately think of any plausible parallelization training scheme which can deliver meaningful updates. Not just every activation, but every parameter is dependent on the training history and updates, like RNNs.
It is possible that training for many serial tokens on non-overlapping datasets could allow for the equivalent of large minibatches, by averaging the final gradients (because we expect parameter updates to be relatively sparse and non-overlapping, while the ‘hidden state’ weights average out). But this is an empirical question.
Using other tricks, like initializing an IFNN from a successful Transformer or AUNN checkpoint, would seem to somewhat defeat the point, and might compromise any of the speculative advantages—if IFNN training produces adversarial robustness, an initialization from a non-robust model might destroy that by providing too strong an initialization to modify substantially and reach a different region in the model landscape.
-
sampling is inherently serial
Unlike an AUNN, which lets you sample arbitrarily, which helps a lot for the weirder use cases like reifying game trees as a giant sparse ‘array’, an IFNN cannot easily jump around. So you need to either much more carefully engineer the encoding to be ‘dense’ in some way or accept an unbounded amount of overhead to increment the implicit pointer through the ‘array’ or come up with some other method still (directly modifying the implicit hidden state, akin to methods like activation steering or lightweight finetuning? Try to encode explicit index jumps as tokenized text commands to let you input ‘increment +152’ and hope that it Just Works™?).
While we can imagine an AUNN being practical someday, the serial limitation on an IFNN is so severe that it’s hard to imagine any improvement could fix it, unless some form of parallel training can be made to work.
Overall, the IFNN is probably most valuable as a thought-experiment or a pedagogical example, to expand our imaginations about what a neural net could be.