How would we estimate the evidence of “Charley Parlapanides”? The names of the writers could either be:
-
present and wrong
Very strong evidence it is fake: who puts their own name down wrong? This would be overwhelming evidence, but we don’t have it so we will drop this possibility from consideration and consider the remaining possibilities:
-
present and right
Evidence it is real. Of the 10 scripts used in the stylometric, 9⁄10 included right authorship information.
-
not present
Of the 4 known fake scripts mentioned previously, only 2 included authorship information.
Given this information, how does the presence of right authorship influence our prior belief of 50%?
Let a be “is real” and b be “has correct authorship”. We want to know the probability of a given the observation “correct authorship”. A version of Bayes’s theorem (stolen from “An Intuitive Explanation of Bayes’s Theorem”; you can see other applications in my modafinil essay; a nice visualization is given by Oscar Bonilla or one could watch distributions be updated):
P(a|b)=P(b|a)⋅P(a)(P(b|a)⋅P(a))+(P(b|¬a)⋅P(¬a))
If you look, the right-hand side of that equation has exactly 4 pieces in its puzzle:
-
P(a)
This is something we already know, “probability of being real”. This is the base rate we already estimated at 50% or 0.5.
-
P(¬a)
This is the negation of the previous. What is the negation of 50%, its contrary? 50%.
-
P(b|a)
Remember, we read the pipe notation backwards, so this is ‘the probability that a real script (a) will include authorship’ (b)’. We said that 9⁄10 of good scripts include authorship, so this is 90% or 0.9. (One way to compensate for the small sample size of 10 scripts would be to use Laplace’s rule of succession, n+1m+2, which would yield 9+110+2=0.83.)
-
P(b|¬a)
Finally, we have “the probability that a fake script will include authorship”. We looked at 4 fake scripts and 2 included authorship, which is another 50% or 0.5.
To put all these definitions in a list:
a = is real
b = has authorship
P(a) = probability of being real = 50% = 0.50
P(¬a) = probability of being not real = 50% = 0.50
P(b|a) = probability a real script will include authorship = 90% = 0.9
P(b|¬a) = probability a fake script will include authorship = 50% = 0.5
We substitute in to the original equation:
P(a|b)=P(b|a)⋅P(a)(P(b|a)⋅P(a))+(P(b|¬a)⋅P(¬a))=0.9⋅0.5(0.9⋅0.5)+(0.5⋅0.5)=0.450.45+0.25=0.450.7=0.643
Sanity checks:
-
Authorship is evidence for it being real; did we increase our confidence that the script is real?
Yes, because 64.3% > 50%. So we moved the right direction.
-
Did we move the right amount?
Well, the fake scripts have a 50% rate and the real scripts have 90%; since this is the only evidence we’ve taken into account so far, our first calculation shouldn’t move us “very far”, whatever that means, since not all real scripts have authorship and plenty of fake ones are careful to include them. (Imagine a world where 80% of fakes include authorship: authorship would become even weaker evidence; and when fakes hit 90% inclusion, authorship would be so weak as to be no evidence at all since the fakes and reals look exactly the same.) The inclusion of authorship does not seem like tremendous evidence so after taking authorship into account, we should be close to our original prior of 50% than to any extreme certainty like 90%.
Are we? Our posterior of 64% doesn’t strike me as a big shift from 50%, so we conclude that this second sanity check is satisfied. Good!
A final calculation: the probability that “a test gives a true positive” divided by “the probability that a test gives a false positive” (P(b|a)P(b|¬a)) is the “likelihood ratio” of that test (see also odds ratio). A likelihood ratio of 1 indicates that our test is useless as it is equally likely for real scripts and fake scripts alike; <1 indicates it is evidence against being real, and >1 evidence for being real. Likelihood ratios will be useful later, so we’ll calculate them too as we go along. So:
P(b|a)P(b|¬a)=0.90.5=1.8
(As expected of evidence for the script being real, the likelihood ratio > 1.)
I also remarked that the use of “Charley” was interesting since there were multiple ways to spell his name. Does this spelling serve as evidence for being real? It turns out: no! It is either irrelevant or evidence against.
To use “Charley” as evidence, we need to know what the real man would be more or less likely to write, and what fakes would be more or less likely to write. I have been unable to find out the “ground truth” here; all 3 variants are used in Google:
“Charles”: 11,800 hits
“Charley”: 182,000 hits
“Charlie”: 1,440 hits
I suspect the truth is likely “Charles” since his Twitter account uses “Charles” (and likewise, Vlas is under Vlasis); his IMDb page lists 5 credits “as Charles Parlapanides” (but nevertheless calls him “Charley”).
What question would we ask here? We could put it as: if we make the assumption that the real man has an even chance of using either “Charles” or “Charlie”/“Charley”, while a fake would choose based on the Google hits (unaware of the variants), how would we change our belief upon observing the script’s use of “Charley”?
a = is real
b = name is spelled “Charley”
P(a) = probability of being real = 64% = 0.64
P(¬a) = probability of being not real = 1 - 0.64 = 0.36
P(b|a) = probability a real script will include “Charley” = 50% (“even chance”) = 0.5
P(b|¬a) = probability a fake script will include “Charley” = 182000182000+11800+1440 = 0.93
Substitute:
P(a|b)=P(b|a)⋅P(a)(P(b|a)⋅P(a))+(P(b|¬a)⋅P(¬a))=0.5⋅0.64(0.5⋅0.64)+(0.93⋅0.36)=0.320.32+0.3348=0.320.6548=0.49
That really hurt the probability, since by assumption using the popular spelling is so heavily correlated with a fake.
Likelihood ratio:
P(b|a)P(b|¬a)=0.50.93=0.538
(We realized the name variant was evidence against, and accordingly, the likelihood ratios < 1.)
Googling “Warner Brothers address” turns up the address used in the PDF as the second hit (it seems to be the official address of all Warner Bros. operations), so we can assume that any faker can find it—if they thought to include it. This question is simply: is a corporate address included? Checking, we see addresses are rare: of the real, 1⁄10; of the fakes, fakes: 0⁄4.
a = is real
b = has address
P(a) = probability of being real = 0.49
P(¬a) = probability of being not real = 1 - 0.49 = 0.51
P(b|a) = probability a real script will include an address = 1⁄10; we apply Laplace’s Rule of Succession to get 1+110+2=212=0.16
P(b|¬a) = probability a fake script will include address = 0⁄4; we apply Laplace (as before) to get 0+14+2 = 1⁄6 = 0.16
Substitute:
P(a|b)=P(b|a)⋅P(a)(P(b|a)⋅P(a))+(P(b|¬a)⋅P(¬a))=0.16⋅0.49(0.16⋅0.49)+(0.16⋅0.51)=0.07840.0784+0.0816=0.07840.16=0.49
0.49? But that was what we started with! It turns out that we are working with such a small sample that when we correct with Laplace’s law, we learn that there are so few instances of screenplays floating around with corporate addresses in them, we can’t actually infer much of anything from it. Does the likelihood ratio agree?
P(b|a)P(b|¬a)=0.160.16=1
(Here we see the final category of likelihood ratios: neither greater than nor less than 1, but equal to 1 - thus neither evidence for nor against.)
We noted the curious fact that while the Parlapanides’ work on the script was announced on 30 April, the PDF claims a date of 9 April.
I did not expect this inversion, but thinking about it in retrospect, this seems consistent with the script being real: the studio commissioned them to write a script, they turned in material, the studio liked it, and the official word went out. (Presumably had the studio disliked it, they would’ve been quietly paid a small sum and a new writer tried.) An ordinary person like me, however, would date any fake version to after the announcement, reasoning that it would be “safe” to date any script to after the announcement.
So we want to express that this inversion is evidence for the script being real, and that frauds would be dated as one would normally expect. If I were to set out to make a fraud, I don’t think I would tinker that way with the PDF date even once out of 20 times, but let’s be very conservative and say a mere 75% of fake scripts would have a normal date (that is: 25% of the time, the faker would be clever enough to invert the dates); and let’s say there was a 50% chance that the real script would be inverted (since we don’t know the real frequency of inversion). The core assumption here is that inversion is more likely for real scripts than fake scripts, an assumption I feel is highly likely (what faker would dare such a blatant inconsistency? It’s Gibbon & the camels again but in a stronger form.) We know how to run the numbers now:
a = is real
b = the date is inverted
P(a) = probability of being real = 0.49
P(¬a) = probability of being not real = 1 - 0.49 = 0.51
P(b|a) = probability a real script will be inverted = 50% = 0.5
P(b|¬a) = probability a fake script will be inverted = 25% = 0.25
Substitute:
P(a|b)=P(b|a)⋅P(a)(P(b|a)⋅P(a))+(P(b|¬a)⋅P(¬a))=0.5⋅0.49(0.5⋅0.49)+(0.25⋅0.51)=0.65772
A jump from 49% to 65.8% is a respectable jump for such a weird date. Then the likelihood ratio is:
P(b|a)P(b|¬a)=0.50.25=2
The metadata date being set in the right timezone is another piece of evidence: a fraud could live pretty much anywhere in the world and his computer will set the PDF to the wrong timezone and he’d have to remember to manually set it to the “right” timezone, while the Parlapanides live in New Jersey and will likely have their PDF timezone set appropriately (even if they travel, as they must, their computers may not go with them, or if the computers go with them, may not change their timezone settings, or if the computers go with them and change their timezone, they may not create the PDF during the trip). So this definitely seems like at least weak evidence.
How to estimate the chance that the fake author would live in a different timezone? If the fraud lived in the US (as is overwhelmingly likely and I’ll assume for the sake of conservatism), the US spans something like 6 distinct timezones. Timezones split up roughly by states so people can estimate the population per timezone; stealing one such estimate:
CST: 85385031
MST: 18715536
PST: 48739504
thus, non-EST: 152840071
EST: 141631478
-
thus, total population: 152840071+141631478=294471549
The US population is more like 312 million than 294 million but the difference isn’t important: what is important is the size of EST compared to the rest of the population.
So, the problem setup becomes:
a = is real
b = is EDT
P(a) = probability of being real = 0.665
P(¬a) = probability of being not real = 1 - 0.665 = 0.3349
P(b|a) = probability a real script will be in EDT = 99% (shit happens) = 0.99
P(b|¬a) = probability a fake script will be in EDT xor the faker will remember to edit the timezone = 141,631,478/294,471,549 xor 0.4 (we assume 0.4 because we used it last time for the PDF creator tool) = 0.481 + 0.4 = 0.881
Substitute:
P(a|b)=P(b|a)⋅P(a)(P(b|a)⋅P(a))+(P(b|¬a)⋅P(¬a))=0.99⋅0.665(0.99⋅0.665)+(0.881⋅0.3349)=0.691
This would have been a much bigger update than 2.6% (from 66.5% to 69.1%) if the evidence of the timezone hadn’t been neutered by our assumption that most fakers would be clever enough to edit it. But anyway, the likelihood ratio:
P(b|a)P(b|¬a)=0.990.881=1.1237
One complicating factor I noticed after writing this section is that Charley Parlapanides’s Twitter page states he lives in Los Angeles, California - not New Jersey. Could they have been living in Los Angeles 2008–200917ya, and the PDF timezone actually be strong evidence against being real? Maybe. My best evidence indicates the move didn’t happen after 201115ya. If the effect of a <200917ya move to Los Angeles were simply to render this argument useless - a likelihood ratio equal to 1 - it would not bother me too much because the likelihood ratio is ‘just’ 1.12, and an error here small compared to errors elsewhere like in the stylometrics analysis. But more realistically, if this argument were wrong, the right argument would likely flip the likelihood ratio to something more like 0.5, and the difference between 1.12 and 0.5 is worth worrying about.
So far so good? No! Vincent Yu points out something interesting: my PDF viewer, Evince, may display timezones as the user’s timezone, not the actual timezone of creation. Is this true? Is Evince misleading me when it gives the timezone as EDT (the timezone I live in)? We appeal to pdftk again: the exact raw date was “D:20090409213247Z”. PHP docs explain the datestamp, particularly the puzzling final character ‘Z’:
CreationDate - string, optional, the date and time the document was created, in the following form: “D:YYYYMMDDHHmmSSOHH’mm’”, where: YYYY is the year. MM is the month. DD is the day (01-31)…The apostrophe character (’) after HH and mm is part of the syntax. All fields after the year are optional. (The prefix D:, although also optional, is strongly recommended.) The default values for MM and DD are both 01; all other numerical fields default to zero values. A plus sign (+) as the value of the O field signifies that local time is later than UT [Universal Time], a minus sign (−) that local time is earlier than UT, and the letter Z that local time is equal to UT. If no UT information is specified, the relationship of the specified time to UT is considered to be unknown. Whether or not the time zone is known, the rest of the date should be specified in local time.
The “Z” says the input date was in UT. Universal Time is a synonym for GMT - so this PDF was created in Europe/England? No; a little more sleuthing turns up the PDF creator software, DynamicPDF, has an API in which the CreationDate is defined to be a java.util.Date object which doesn’t deal with timezones but instead defaults to UT/GMT. So, the timezone doesn’t exist in the metadata; it never existed; and it never could exist in data produced by this PDF creator software.
We could try to rescue the timezone argument by shifting the argument to pointing out that the PDF creator software could have been a type which correctly stored the original timezone in the metadata, which could then provide evidence against being real if the timezone were not EDT, so we could regard this as a very weak piece of evidence in favor of being real - a possible counterpoint turned out to not exist - but this is now so tenuous it is better to drop the argument entirely.
We could isolate multiple tests here from my freeform observations:
-
length
Some of the fake scripts are very long and complete; I remarked in an earlier footnote that the fake Batman script is actually too long for a movie. One of the fake scripts was a single leaked page, making for a 3⁄4 rate.
-
formatting
The sample of real scripts has been reformatted for Internet distribution and doesn’t include the “original” PDFs or representations thereof; worse, the 4 or 5 fake scripts are all properly formatted. With the existing corpus, this test turns out to be useless!
With the dubious benefit of hindsight, we might claim this is not a surprise: after all, any script without formatting would be “obviously” a fake and one would never hear about it. One only hears about plausible fakes which possess at least the basic surface features of a real script.
-
writing quality (spelling & grammar)
In addition, the fake scripts are well-written. Like formatting, this turns out to be a bad indicator; someone writing a movie-length script seems to also be the sort of person who can write well. The description of one of the fakes is interesting in this regard:
This is probably one of the most elaborate ruses on the list. The script was written by 27-year-old Los Angeles writer Justin Becker, and as far as we can tell, he did it for laughs. Becker traveled across the West Coast, planting his scripts all over bookstores, hoping they would get discovered. He basically thought, “it would be funny to find out that a Mr. Peepers movie had been written, and it was very serious and pretentious and political, and it had been shelved because of 9/11” (SF Weekly), which is explained in the preface of the script and by the fact that the screenplay was supposedly written one day before September 11th, 200125ya and contained George W. Bush in the story.
This leaves just length as a test:
a = is real
b = is full-length
P(a) = probability of being real = 0.691
P(¬a) = probability of being not real = 1 - 0.691 = 0.334
P(b|a) = probability a real script will be full-length = 99% (shit happens) = 0.99
P(b|¬a) = probability a fake script will be full-length = 3⁄4, by Laplace, 3+14+2=46 = 0.66
Substitute:
P(a|b)=P(b|a)⋅P(a)(P(b|a)⋅P(a))+(P(b|¬a)⋅P(¬a))=0.99⋅0.6650.99⋅0.665+0.66⋅0.334=0.749
Likelihood ratio:
P(b|a)P(b|¬a)=0.990.66=1.5
The earlier plot summary conveyed the “Hollywood” feel of the plot but unfortunately it’s hard to judge from localization: a DN fan attempting to imitate a Hollywood-targeted script might rename Light to “Luke”, might simplify the plot considerably (there is precedent in the Japanese live-actions movies Death Note, Death Note: The Last Name & L: Change the World), might set it in NYC (Tokyo is out of the question, as Hollywood movies are never set overseas unless the plot calls for it specifically, and NYC seems to be the default location of crime-related movies & TV shows), and so on.
Some of the plot changes make more sense after reading the biography of the Parlapanides brothers: they are Greek and live in New Jersey. Changing “Light” to “Luke” is a very clever touch in localizing the character: besides the visual resemblance of being short one-syllable names starting with “L”, apparently “Luke” is a form of “Lucius”, better known as “Lucifer”, and the Latin was literally “light”! (And indeed, Luke seems to still be a common Greek name, perhaps thanks to the Gospel of Luke). NYC is a the default location, but it’s even more natural when you are 2 screenwriters who grew up and live in New Jersey. (I grew up on Long Island, and for me too, NYC is simply “the city”.)
More importantly, the plot includes several idiot-ball-related changes that I think any DN fan competent enough to write this fake would never have made, even in the name of localization and Hollywoodization: the incompetent bus ID trick comes to mind.
Unfortunately, in both respects, I can’t assign defensible numbers to my interpretation for the simple reason that any reasonable differences in probabilities leads to a ridiculously strong conclusion!
For example, if I gave 90% (fakes) vs 95% (real) for the individual localization points (for each of name, simplification, location), and then 25% (fakes) vs 50% (real) for 2 instances of incompetence, this gives us a likelihood ratio of:
0.950.90⋅0.950.90⋅0.950.90⋅0.500.25⋅0.500.25=4.7
(Here we see an advantage of likelihood ratios: they’re easy to calculate and give us an indicator of argument strength without having to run through 5 different iterations of Bayes’s theorem! This is something one learns to appreciate after a few calculations.)
A likelihood ratio of 4.7 would be the single strongest set of arguments we have seen yet, and even stronger than the stylometric likelihood ratio in the next section. If we used this result, it would be solely responsible for a very large amount of the conclusion. A critic of the final conclusion would be right to wonder if the conclusion rested solely on this dubious and unusually subjective section, so we will omit it (with the understanding that as usual, we are being conservative and essentially trying to calculate a lower bound to compensate for arrogance or overly favorable assumptions elsewhere).
The stylometric result is straightforward: if a fake script gets paired up randomly, then it had just a 1⁄15 chance of pairing up with Immortals. Even if we restrict the matches to the other movie scripts, there were 10 movie scripts and 2 oddballs for 12 total or 6 pairings, giving 1⁄6 chance of randomly pairing up with Immortals. The real question is: if the script is real, what chance does it have of pairing up with something else by the same authors? I included 4 fanfictions by the same author (Eliezer Yudkowsky), and 2 wound up pairing (with the other 2 in the same overall cluster but more distant from the pair and each other), giving a rough guess of 50%; this is convenient since our default “I have no idea at all” guess for any binary question is 50%, and even if we apply Laplace, we still get 50% (2+14+2=36 = 50%). So as usual, we will make the most conservative assumption for the fake, and keep our pessimistic assumption about the real.
a = is real
b = is paired with Immortals
P(a) = probability of being real = 0.749
P(¬a) = probability of being not real = 1 - 0.7703 = 0.251
P(b|a) = probability a real script will be paired with Immortals = 50% = 0.50
P(b|¬a) = probability a fake script will paired with Immortals = 1⁄6 = 0.1667
P(a|b)=P(b|a)⋅P(a)(P(b|a)⋅P(a))+(P(b|¬a)⋅P(¬a))=0.50⋅0.749(0.50⋅0.749)+(0.1667⋅0.251)=0.899
P(b|a)P(b|¬a)=0.500.1667=2.999
As expected, the stylometrics was powerful evidence.