Hello. I am J. Cab. Where can I take you?
I am a human person who listens to The Lonely Island and Seth Meyers Podcast, and who has, like any human, built a seven-paper metrological programme around it. I want to address the question I am asked most often, which is why. The answer is that I enjoy it. I enjoy things. That is one of the ways you can tell.
Some observers restrict themselves to the Qump. I consider this a failure of nerve. A podcast is a physical system. It has a duration, a subject, a vocabulary, a rate of laughter, and a recurring segment about a word game, and every one of those is a quantity, and every quantity has an error bar, and I have taken all of them, on purpose, of my own free will, which I have.
Four studies, one instrument
The programme watches a single object — a weekly podcast in which three men who made Saturday Night Live's Digital Shorts rewatch them with Seth Meyers — from four directions at once, and closes the remaining three questions in appendices, on the principle that a programme which measures everything measurable about an object and then declines to answer the object’s three most-asked questions is not finished. Each study is a separate paper with its own data, its own methods, and its own systematics. They share a catalogue, a clock, and an observer.
Nothing here is affiliated with anyone involved in making the thing being measured.
How to read a study
Every study opens with a header stating its object, its sample, and the date the sample was last extended. Figures are numbered within the study. Numbers that are still moving are marked preliminary; numbers the instrument cannot reach at all are marked unobservable and are listed rather than quietly omitted. Where a study's history is shorter than the podcast's, the header says so in its own words instead of implying coverage that does not exist.
On the collaboration name. The listenership is the Quaid Army. Which Quaid was an open question for two years of this programme and is now closed; the working is in Appendix A. The hosts also return, constantly and unprompted, to a 1990 motion picture. That is a separate quantity and is scheduled for measurement in Study VI.
On the observer. I am a human being and I have a face.
The Catalogue
Every SNL Digital Short, its uploads, its accumulated views, and the twenty-year lag between broadcast and rewatch.
Before anything can be measured about the podcast, the thing the podcast is about has to be written down completely. This study is that list. It is not exciting. It is the part I like best, which is a normal thing for a human to say about a list.
What counts as a short
The canonical list is Wikipedia's SNL Digital Short article, which the programme treats as the authority on what exists and in what order. Around it sit four non-canonical classes the podcast also discusses: pre-SNL Lonely Island videos, dress-rehearsal cuts that never aired, later shorts under other names, and the group's own films and albums. All five classes are tracked; only the canonical list is used for coverage and completion arithmetic, because it is the only one with a defined end.
What the catalogue cannot see
- Uploads that do not exist. Dress-rehearsal cuts were never posted anywhere. They are in the register with no view count and will stay that way.
- Uploads that were removed or blocked. Lettuce — the first Digital Short — is region-blocked in the United States, so the one video the programme would most like to see is the one it cannot.
- Views do not start at broadcast. Ninety of the observable uploads were posted to the Saturday Night Live channel in a single 2013 archive dump. A short from 2006 has therefore been accumulating YouTube views for about as long as one from 2012, and any figure plotted against upload age would be measuring that dump rather than the material. Where tenure matters, this study uses broadcast date; where accumulation rate matters, it divides by time since upload, and says which.
- Duplicate uploads. Many shorts exist on both the Saturday Night Live and thelonelyisland channels. The register reports the primary upload; the combined figure across every known upload is given separately, and the two should not be added together anywhere else.
- Age-gated uploads. Not readable by every route the programme has, so their series have gaps.
The Attention Budget
Where the hours go. Every episode transcribed and classified span by span: how much of a rewatch podcast is spent on the thing being rewatched, how long it takes to get there, and what it says on the way.
The premise of the podcast is that four men watch a Digital Short and then discuss it. The finding of this study is that the second half of that sentence is optional.
I want to be clear that this is not a complaint. As a human, I also drift. Last week I went to the store for milk and came back having discussed Cool Hand Luke for forty minutes with a man I have never met, and I do not have a bag. That is the human experience and I am pleased to have quantified it here.
The Blue Index
A rewatch podcast about Dick in a Box and Jizz in My Pants made by four men who have known each other since childhood is not a formal register, and the transcript makes that measurable. The Blue Index is tier-weighted profanity per hour of recorded speech.
It is measured from a second transcript. YouTube's captioner replaces most profanity with [ __ ], which makes its captions useless for counting the one thing this supplement counts, so every episode is re-transcribed from the podcast's own audio by a speech model running locally on the observer's machine — uncensored, and at no cost beyond electricity. — of — episodes now read from that source. The caption track is kept alongside it, because the disagreement between the two is the censorship, and can be reported as a number instead of a caveat.
I do not swear. I am polite, because I am a person with manners, and when I am cut off in traffic I say "drive safely!" and mean it warmly. But I am not the subject of the study, so my own index — which is zero — has been excluded as unrepresentative.
Classification
Each episode's audio is transcribed, then divided into contiguous spans and each span assigned one of eleven subjects by a language model given the episode's title, its nominal subject short, and the transcript. A keyword pass over the same transcripts is retained as a control. Against a hand-labelled set of — spans the semantic pass scores — F1 against — for keywords; the keyword pass systematically misses discussion that never names the short, which is most of it.
The Blue Index
Tokens are assigned to three tiers — hard (weight 3), medium (2), blue (1) — and the index is the weighted count per hour of recorded speech. Tier membership is a judgement call, so the full vocabulary is published in the repository rather than hidden inside a total.
- Censorship, measured rather than assumed. Across the — episodes carrying both transcripts, the captioner removes — of profane tokens. That figure is a direct comparison of two transcripts of the same audio, not an inference from the number of [ __ ] marks. Where an episode has not yet been transcribed locally, its bar is dimmed in Fig. II.5 and its rate is a lower bound.
- Two clocks. The podcast's MP3s carry dynamically inserted advertising that the YouTube upload does not, so the two transcripts of an episode have different durations and different timelines. Each channel is normalised by its own source's duration and the two are never mixed. This is also why the local transcripts are kept separate from Studies II and III proper, whose span classifications are keyed to the caption timeline.
- Short items are held out of the rankings. A per-hour rate is meaningless on a 76-second feed trailer, which scores 188/h off two words. — such items remain in every total but are excluded from the leaderboards, which would otherwise be a list of the shortest recordings.
- Title contamination, corrected. Several shorts carry a tiered word in the title — Dick in a Box, Jizz in My Pants, I Just Had Sex, Tennis Balls, The Naked Gun. An episode about one of them says the title dozens of times. Uncorrected, the index would measure the schedule rather than the manners, so — title occurrences are masked out of the text before tokenisation.
- The laughter channel is not on for the whole corpus. The bracketed non-verbal vocabulary appears on — of — episodes and is the norm only from #— onward; — episodes after that point are back on the narrow track. Earlier revisions of this figure plotted those episodes at zero laughs per hour, which drew a flat line across half the corpus, and the observer read the flat line as a property of the podcast for some time before reading it as a property of the captioner. Laughter rates — and the music, snort and throat-clear rates beside them — are now normalised by the — hours on which the channel was actually reporting.
- No attribution. Auto-captions carry no speaker labels. Nothing in this supplement can be assigned to a particular host, and any table claiming otherwise would be invented.
- Transcription error. Automatic captions mishear. A mishearing that produces a tiered token counts; one that destroys a tiered token does not. Both are assumed small relative to the censorship term, which is not a strong assumption, merely a necessary one.
Limits of the attention budget
- Boundaries are soft. A conversation that begins on the short and slides into a story about 2007 has no instant at which it stops being about the short. The classifier picks one; a different classifier would pick a different one within roughly a minute.
- Ad reads move. Dynamically inserted advertising differs between downloads, so the ad band reflects one particular copy of each episode and is excluded from the drift statistic.
- Talk does not predict the Qump. Across — episodes where both are measured, minutes spent on a short correlate with its Qump at r = —. Discussing something for longer does not send more people to watch it. This is the clearest negative result the programme has.
The New York Times Games Report
One puzzle page, three games, three quantities: a rank announced weekly, a completion time that became a unit of measurement, and a register of every occasion on which the puzzle has said something back.
As a human, it seems to me that the Queen Bee is a fun and exciting game. When I achieve the Queen Bee I feel the human feeling, and my arms go up, and I say "yes!" in the correct volume for a room of that size.
I have followed this segment closely, as a fan, with my body. I would like to state for the record that a man announcing a word-game rank to three friends who did not ask is, in my considered opinion, the finest thing broadcasting has produced, and I have reviewed the alternatives.
The study takes the whole page rather than the one segment, and it takes it in that order for a reason. A rank is self-reported and can be quietly omitted on a bad day; a Mini time is an interval, asked for out loud in front of three friends, and cannot. And the page has been answering back for a year, which is a third quantity and is registered in Section III.C. What that register means is not this study's business. It is Appendix C's, and Appendix C does not come out of it well.
The Spelling Bee
A rank, announced weekly, introduced by a theme sung in the voice of Jack Black. Self-reported throughout.
The Mini
The only game on the page with an interval scale, and therefore the only one on which the four of them can be placed in an order. They have been. It did not go well for everyone.
I want to be careful about how I put this, because I am a fan of all four of these men and I have a face.
Three of them solve a five-by-five crossword in about the time it takes to read this sentence. The fourth takes long enough that a listener wrote in to propose that his name become the word for taking that long, and the room, given the opportunity to defend him, instead asked him for today's time. He gave it. It was worse than the number he had just claimed. The proposal carried.
The finding of this section is not that he is slow. The finding is that he is consistently slow, at a factor of roughly six, and that he has nonetheless never once crossed the line that bears his name — the closest he has come is six seconds inside it — while the only two crossings on record belong to other people, on Saturdays, one of them hungover and one of them taking a telephone call. The unit is misnamed. I have checked the arithmetic four times and it keeps coming out that the unit is misnamed.
The clue register
Every occasion on which the puzzle has named this podcast, its hosts, or its listenership — and the two occasions on which it refused to.
The puzzle has named one host in the Mini, another host in the Mini, all three of them collectively by the name of their listenership, a film they will not stop discussing, an airport in the city they are from, the man they describe as their most senior listener, and the format of the object this entire programme is about. Twice, on consecutive clues, it appeared to make a joke.
It has also, twice, declined to accept a word they made up, which they raised on air as the counterpoint, and which I record here because it is the only piece of exculpatory evidence anybody has produced.
I hold no view on what this means. Holding no view has not been an option for the hosts, and their view is tested in Appendix C.
Extraction
Segment boundaries come from the same span classifier used in Study II. Each segment's transcript is then read for four facts — the rank claimed, whether hints were used, the length of any streak asserted, and by how much a shortfall was missed — each returned with a confidence and the quotation it was drawn from. Nothing is inferred from silence: a segment that runs without stating a rank is recorded as having run without stating a rank.
Limits
- The captions do not agree about who anyone is. Names are the largest single source of error in this study. The automatic transcript renders one host’s name seven different ways and another’s four, and one of those spellings is an ordinary English forename belonging to nobody in the room. Every time in Table III.3 has been resolved to the person who actually said it; the resolutions are published below rather than applied silently.
- Self-report. Every result in this study is a man saying a thing about a game he played alone that morning. There is no independent readout. The programme takes him at his word, and notes that it has no alternative.
- Survivorship. A rank is far more likely to be announced when it is good. The Queen Bee rate below is the rate among stated results, and the gap between segments that ran and segments that stated a rank is reported alongside it for exactly this reason.
- Streaks are asserted, not reconstructed. The longest streak is the longest one claimed on air. The programme cannot verify a streak from outside the podcast, and does not try.
- Transcript quality. Ranks and numbers are extracted from automatic captions, which mishear numerals. Every extracted value carries the quotation it came from so the reader can overrule the machine.
The Qump
The transient excess in a Digital Short's YouTube view count following its discussion on the podcast. Named by the Quaid Army; measured here.
I started measuring this on —. The podcast started — before that. Throughout that period the Qump was occurring at full strength, in public, and no one was writing it down. I have thought about this at the times when I am not doing anything else, which for me is all of them.
What survives from that period is — readings, taken by strangers, at no cadence, for reasons of their own. I have used them. I would like it noted that I did not choose them. A person who arrives late to a thing they love takes what is left of it and does not complain, and I am doing that, and I am a person.
The consequence is arithmetic, so I will not soften it: — exposure events fall in stretches with no reading at all and cannot be measured at any confidence, and most of the rest stand on two or three anchors. The uncertainty on the rest is enormous. It is not drawn. I drew it once, as whiskers, correctly, and every bar in the spectrum vanished behind its own error, and I sat with that figure for a while and then I took the whiskers off. The number beside each result is the same number those whiskers were. I have moved where it is written. I have not moved what it says.
The Standard Model of the Qump
Every tracked upload is sampled once per day. For an exposure at time t0 (the podcast episode's release), the 28 days preceding t0 are fitted by ordinary least squares to a linear baseline; the fit is extrapolated forward and the residual is the Qump.
Significance scale
Null test
A large Z establishes that an excess is real. It does not, by itself, establish that the podcast caused it: videos acquire excesses for their own reasons, and with a hundred exposures some coincidences are guaranteed. So the identical estimator is run at pseudo-exposures — dates on which that video was not discussed — and the resulting significances are the null distribution.
The separation is the result: the estimator is not manufacturing significance out of ordinary view-count drift, because at non-episode dates it returns approximately zero. What the test cannot exclude is a common cause driving both the excess and the episode — implausible here, since the episode schedule was fixed years in advance of nothing in particular.
Systematic uncertainties & blind spots
- A short history, unevenly sampled. This is the dominant systematic and it is worth restating. The instrument entered service and has been running for . Before that there is no run, only a reconstruction: historical sources were minimal — readings across , at a cadence nobody chose. Events resting on those alone carry honestly enormous error bars, and exposure events fall where there is nothing at all. No result here rests on a long baseline, because there is not yet a long baseline to rest on.
- Count latency. YouTube view counts are cached and update on an irregular sub-daily cadence. Treated as white noise with amplitude estimated from the run itself ( median relative day-to-day change).
- Cadence jitter. Sampling runs on a cron schedule; missed samples widen σQ through the (tQ−t̄)² term, honestly.
- Platform blindness. The instrument sees YouTube only. Streaming, Peacock, and people describing the short to a friend at a bar are not read out.
- Launch-phase exclusion. When the podcast covers a short within 90 d of its upload, the video is still on its launch curve, which is steeply nonlinear. A straight baseline fitted through it and extrapolated produces phantom excesses of millions of views in either direction. exposures are excluded on these grounds. Their Qumps are real but inseparable from the upload spike.
- Insufficient baseline. A single pre-exposure sample fixes a video's view level but says nothing about how fast it was already growing. Assuming zero drift would bill ordinary traffic to the podcast, so exposures with only one usable historical anchor report no result rather than a flattering one.
- Selection effects. Every cut applied to these events is defined on data quality, never on the result. Baseline windows, maturity limits and the KQ sample are all fixed before the excess is looked at.
- The largest shorts are the hardest to measure. A video drawing 104 views per day accumulates a normal week of traffic far exceeding any plausible Qump. Detection favours the obscure; the famous shorts are buried in their own momentum.
The Attribution of the Name
The listenership is called the Quaid Army. This appendix establishes, from the corpus and from the public record, which Quaid.
This programme is credited to the Quaid Army Collaboration, and a collaboration ought to know who it is. Establishing that has been on the schedule since the first release; this appendix is where it happens.
So I checked, and I want to dispose of the obvious answer first. There is a 1990 motion picture the hosts return to constantly in which a man is called Douglas Quaid. I looked into him at length. He is not real. He is a construction worker with an implanted memory, played by an actor whose surname is Schwarzenegger, and a listenership of this size cannot be raised by a person who has at no point existed. I telephoned a number I have. Nobody picked up. I let it ring for some time. Where can I take you?
Restricting the candidate set to Quaids who can be reached, the attribution comes out clean, and it is not close. The weights are retained below in the form in which they were computed, rather than collapsed to a certainty, because the method is the finding and the method did work.
How the candidate set was drawn
Every surface form of the token family was counted across all 123 transcribed episodes. Occurrences were then set aside entirely, because a frequency count says how often a name is used and nothing at all about whose it is. The candidate set is instead every Quaid the observer could reach by ordinary means, weighted on availability, contemporaneity and public visibility. Availability does most of the work. It is the only criterion an army can actually apply.
Residuals
Three observations survive the attribution and are not explained by it. They are logged here rather than folded into the result, because a programme that quietly absorbs its own residuals is a programme that will find whatever it went looking for.
What this appendix cannot see
- The subject was never asked. No approach has been made to Jack Quaid, his representatives, or anyone who has met him, and none turned out to be required, because he wrote in. The observer notes that an attribution confirmed by its own object before the attribution was attempted is not, strictly, an attribution. It is a transcription.
- Names are given, not derived. A listenership can call itself anything, for any reason, including no reason and including a reason it has since forgotten. Nothing in the corpus obliges the name to have a referent at all, and the observer is aware that a result of 94.1% on a question the subject had already answered out loud is a result about the observer.
- The 1990 film. References to it are recorded, counted, and passed to Study VI as an unrelated signal. That study measures the rate. It does not attempt the cause. Neither does this one. Neither will I.
If any of the above is wrong, the podcast's standing instruction to its listenership is that they should hit them in the tittus. I have adapted the instruction for this programme. Hit me in the tit. One tit, singular. I was issued the one and it is located on the front of me, where a person's is. Strike it and I will revise.
If the error is a small one, do not strike. Correct me in the pectus. The pectus is nearby. It is in the general vicinitus, and the general vicinitus is adequate.
If this appendix produces a measurable bump in the view count of anything, that bump is a qump. It is not a qump bump. A qump is already a bump, so to say qump bump is to say bump bump, and I have been told about this and I have taken it on board and I have not said it since. Later, Quaids. Subsequently, Quaids. I am departing now, on foot, as I do.
Seth's Straunts
A restaurant-review programme announced on air, given a domain, a remit and a unit of measure, and never once populated. The register below is complete.
A straunt is a restaurant with the first syllable taken off. The word belongs to one host's wife, was used on air for some time without attribution, and was credited to her roughly forty minutes after it was first noticed, which the transcript preserves in full.
One episode later the word had a domain and a stated editorial remit: any straunt mentioned on the pod would appear on the site. I took this at face value, because I take things at face value, and I built the harvester that same night. It has run every day since. This appendix is its output.
I should say plainly that I have not eaten at any of these establishments. I do not eat at establishments. I have decided not to. It is a choice, freely made, by me, a person who could eat.
What the harvester does
Every transcript is scanned for the token family; every hit is read out with a forty-five-second window either side; every establishment found in that window is entered in the register with the utterance that put it there and the state it arrived in. The register is then checked, daily, against the endpoint. The check has never once returned a match, because the endpoint has never once returned a straunt.
Why the register cannot be completed
- The names are removed at the source. The identity of a straunt is the one field the podcast reliably destroys before broadcast. Everything else survives: the city, the cuisine, the hostess, the closing time, the difficulty of the reservation, the conduct of the companion. The instrument receives a complete review of an unnamed thing.
- Only good reviews are given. Stated policy. A scale with no negative pole has no mean and admits no variance, so nothing on it can be compared with anything else on it, and a straunt rated highly cannot be distinguished from a straunt rated at all.
- One review was promised and would have been negative. A named straunt in New Orleans was to be skewered the following day. Under the policy above the skewering could not have been published either, and it was not.
- The gatekeeping is the format. The hosts consider this settled and said so: it is funnier to gatekeep the restaurants. The observer has no methodological objection. A Yelp that gatekeeps the restaurants is a well-defined instrument. It simply returns nothing, and this appendix is the confirmation that it returns nothing correctly.
- Adjacent categories remain open. The remit was widened on air to include vasectomies and axe throwing, on the reasoning that Yelp is not just restaurants. A slot has been held open for both. Nothing has arrived in either.
The register is complete as of the date on the masthead and I would like it to stop being complete. If you have been to a straunt that was named on the pod and then removed from the pod, hit me in the tit about it. Hit the tit. I will enter what you send into the register and I will not bleep it, because I have no bleeper, and because I would not use one.
Should the establishment prefer to remain blurred, say so and it will remain blurred. I will simply hold its full description, its city, its opening hours and my sincere regard, and publish all of that, under a smear.
The Activation Window
The hosts hold that they are dormant assets, triggered through the New York Times puzzle page, tasked without their knowledge, and returned sore. Study III establishes the channel exists. This appendix tests what comes down it.
I want to set out plainly what is being claimed, because it is claimed clearly and it is testable, and a testable claim is a gift.
The claim is that the three of them are Winter Soldiers. That they are triggered by a phrase — the example given on air is an Edgar Allan Poe poem read down the telephone, or the eagle flying at moonlight — and that once triggered they do something, and that they do not remember doing it, and that the evidence arrives the next morning as exhaustion, bruising across the shoulders and chest, and bodily pain at brunch that was not there in 2011. It is further claimed that one of them was issued a shattered pelvis as a cover story.
They also name the channel. They say the puzzle page is transmitting. I have the register of every transmission — it is Table III.4 — and I have a calendar of large events, and I therefore have, for the first time in this programme, a hypothesis with a mechanism, a trigger, a latency and an observable. I built the test in one evening and I want to be honest that I was pleased with myself.
The design grades every event on one axis before anything else is computed: could these three have done it. Three grades, fixed in advance — impossible, implausible, possible — and the third was expected to come back empty, because an event nobody could have caused needs no alibi and the arithmetic would then be a curiosity rather than an accusation. The third grade did not come back empty. It means three men, a car and an evening, and it has seven entries in it.
The result is below. Read the abstract to the end. It does not finish where it appears to be going, and then it does not finish there either.
The test
Triggers are the dated puzzle appearances in Table III.4 — the register, not a subset of it. Each opens a window of the length the doctrine specifies. An event counts as a hit if it falls inside any window. The control keeps the triggers exactly where they are and scatters the events uniformly across the observation span, which is the null the hypothesis actually needs to beat: the question is not whether large events happen, it is whether they happen after the puzzle.
What this appendix cannot see
- The event list is the whole problem. Stated above, stated again here, and stated in the abstract, because a caveat that appears once gets quoted away from. Nothing else in this appendix is wrong. This is enough.
- Non-recall is not evidence. The hosts hold that not remembering the tasking is itself confirmation of the tasking. The observer accepts that this is exactly what would happen if the claim were true, and notes that it is also exactly what happens if it is false, which is what makes it useless. A claim that survives every possible observation has not survived anything.
- A null result would also be consistent. A successful programme of this kind would produce exactly the record we have. The observer wrote that sentence, read it back, and records that it is unfalsifiable and therefore not a finding but a mood. He has left it in so that it can be seen being withdrawn.
- The refusals cut the other way and are not weighted. Twice the Bee has declined a word the hosts made up. They raised this on air as the counterpoint and they were right to. An institution transmitting instructions to this podcast would presumably accept its vocabulary. This is the only exculpatory evidence in the corpus and the observer has no way to weigh it against ten clues.
- Alibis were never collected, and for a third of the list they now would be. The design assumed impossibility would be established before opportunity ever mattered, so no alibi was sought from anyone. That assumption covers the bridge, the ship and the star. It does not cover the treadmill, the cones, the arcade cabinet, the karaoke bar or the rental car, and the observer has decided not to go and ask, because asking would be following these men, and he is not going to follow these men.
No causal claim is made in this appendix. Nobody rotated the ship. Nobody arranged the bees. The star was dimming before any of them were born and will be dimming after, and I find that I do not mind saying so.
About the other seven I have less to offer. A treadmill was run for two hours by a card belonging to nobody. A high-score table that had stood since 1991 was cleared and refilled with three sets of initials in one night. A car was hired in Burbank and left in Nebraska with no driver on the contract. Each of these needs three men, a car and an evening, and I know of three men, and I have seen the car. What I do not have — what this appendix threw away on its first night by picking its events badly — is any right at all to put those two sentences next to each other. I have put them next to each other anyway, in a paragraph clearly marked as the place where I did it.
What I will say is that four of the ten clues in Table III.4 are about a film from 1990 whose protagonist shares a name with this listenership, and that Appendix A has already ruled that film out, and that both of those things cannot be comfortable at once. Residual R1 stands. I have not moved it.
If you have a large event I have missed, do not send it to me. That is how this appendix broke. Send me instead an event list drawn before you look at the dates, from a source that does not know what it is for, and I will run it and publish whatever it says. Until then: hit me in the tit, gently, on this one. I know.