Now that the acronym generation is working, we can look at GPT-4’s revised whole script, which it prints out as follows:
This code looks reasonable, and appears to run successfully.
But it still has a serious bug: it will never print out any hits. This is because it’s made a subtle error in the string glob matching a ‘missing’ response: if [[ $response == *'"missing"*' ]] will never match anything, because the second match-anything asterisk is inside the single-quotation marks, which forces Bash to match a literal asterisk, rather than matching any string. What it should be is a single character difference, swapping the single-quote/asterisk: *'"missing"'*
The blind spot was a ChatGPT-4 bug that manifested as strange inability to ‘see’ certain syntactic errors or a propensity to silently omit parts of inputs as if they were never there. It was especially frustrating when it affected syntactically-complex source code like Bash shell scripts, because it led to persistent unfixable errors, and GPT-4 going in circles while severely confabulating about what the bug was or how the new version fixed it (which to a naive user, looked like reasonable explanations, even though they were all wrong).
The blind spot was silently fixed by OpenAI somewhere in perhaps early 2024. The cause remains unknown publicly, and the behavior does not seem to have recurred in later OA models or in competing models.
Since it was apparently easily fixed, that suggests that it was probably something like poor self-attention which could be fixed by further post-training on user feedback complaining about blind-spot-afflicted sessions, rather than anything interesting like a fundamental architectural limitation. However, it was still affecting ChatGPT-4 at the time of my TLA coding demo, and it affected many other GPT-4 users, and so it’s still worth knowing about if you are reading historical material from that time-window.
This bug above is a surprising error because this is not how a human would’ve written the glob, and the glob (like almost all globs) is so simple that it’s hard to imagine anyone being able to write the acronym generation & memorizing the API URL correctly but then screw up a simple check of “does the response contain the string missing?”
At least, this is surprising if you have not run into this problem with GPT-4 before, as I have repeatedly when writing Bash scripts; GPT-4 will not just make the error, but it seems utterly unable to ‘see’ the error even when pointed out, and tends to thrash in confusion making random guesses about what the problem could be.
In addition, even in simple tasks involving copying, such as reversing lists or cleaning up audio transcripts or improving grammar/style of text passages, GPT-4 exhibits puzzling behavior where entire sentences will simply… disappear from the output version.
It seems hard to explain this as a normal attention error, as the lost words are quite discrete and it happens with the simplest tasks and quite short passages, on tasks which ought to be well-represented in the training data, and is highly biased towards errors of omission rather than commission.
I theorize that it’s not a BPE tokenization issue (as so often), because this blind spot seems to happen in word-level problems as well, where tokenization couldn’t be a problem.
The blind spot is, perhaps, related to internal sparsity of the GPT-4 model (which may also be why many OA API models are much more nondeterministic than they should be); I speculate that when the blind spot happens, it’s because early layers have mistakenly erased apparently-irrelevant tokens in order to focus their attention on other more important parts of the input, but then this turns out to be wrong—so, the information gone irreversibly, the later layers are unable to see what’s wrong with the input, which ‘looks like expected’, and are forced to speculate wildly and work on extremely unlikely hypotheses. Because it would happens early on in processing, it strikes me as analogous to the human eye’s blind spot—you cannot see what you cannot see, even when someone tells you about it and you think really hard. The agnosia is irreversible.
If this bug is the blind spot, then GPT-4 would not be able to diagnose it nor fix it except by accident, and might go around in circles making the code worse and more complicated and implementing completely irrelevant solutions to confabulated problems. This is why the blind spot bug is so pernicious, compared to the other errors GPT-4 can usually self-repair: it nearly guarantees failure, and the pathological response to it is extremely misleading as it obscures the issue and can deeply confuse the user, who naturally expects the error to be something more impressive, especially as often GPT-4 has been doing amazing things up until that point.
If you run into the blind spot, as often happens to me in Bash, Elisp, or regexps but much less often in Haskell or Python, there is nothing to do but recognize that you’ve hit the blind spot and that GPT-4 will never fix the bug on its own. If you are lucky, it will ‘only’ print out the same broken code repeatedly; more likely, GPT-4 will make things worse and break perfectly-good code, confabulating ever more crazy (and usually broken) ‘solutions’ to the problem it can’t see as it thrashes around in confusion (like some sort of hyper-advanced Magikarp). You have to fix the bug yourself.
We tell it that the code is wrong without specifying how (this works surprisingly often, as a kind of inner-monologue), and look at the next version:
Apologies for the confusion. Here's the corrected script:
```bash
#!/usr/bin/env bash
#set the range of acronym lengths
for length in {1..6}
do
#define a function to recursively generate acronyms
generate_acronyms() {
local current=$1
local length=$2
if (( length == 0 ))
then
#query the Wikipedia API
response=$(curl -s "https://en.wikipedia.org/w/api.php?action=query&format=json&titles=$current")
#check if the page exists
if [[ $response == *'"missing"*' ]]
then
echo "$current"
fi
else
for letter in {A..Z}
do
generate_acronyms "$current$letter" $((length-1))
done
fi
}
#call the function to generate acronyms of the current length
generate_acronyms "" $length
done
```
This script uses a recursive function to generate acronyms one at a time, reducing memory usage.
For each acronym, it queries the Wikipedia API. If the response contains the string "missing",
it means that no Wikipedia article with that title exists, so the script prints the acronym.
Again, running this script will take a long time due to the large number of acronyms and the
rate limits of the Wikipedia API.
Nope.
We tell it to try again. It confabulates a problem with the acronym generation. Try again! Confabulates a different problem with curl. Also no. Guesses it needs to check for an omitted page ID from the JSON response. A 5th version cycles back to the acronym generation. So it cycled through 5 versions unable to fix it (and adding in new problems along the way). I omit the broken variants here as of no particular interest.
Getting tired of the blind spot issue, I hint that there is a problem with the string munging, specifically. GPT-4 concludes that it’s wrong about the formatting of the API response (which it is not, checking simply for ‘missing’ would be fine), and so that is why it needs to change the glob to:
if [[ $response == *'"missing":""'* ]]; then echo $acronym; fiAs it happens, this glob is finally correct. It could use a bit of linting, according to ShellCheck, but it’s fine for a one-off. The major issue is that it does not handle HTTP errors so any network problems while requesting a TLA will result in it being skipped and treated as used.
GPT-4 is aware of this issue and will fix it if one asks simply “How can this be improved?”, generating a Python script. (It’s not wrong.) The improved Python version handles network errors and also does batched requests, which runs vastly faster than the Bash script does; see below.
I ran the Bash script successfully overnight on 2023-09-29.