Every owner-operator, every commission salesperson and every contractor faces the same daily question: today is going well, do I keep going or do I stop. A famous piece of research said people stop, which would be an error. Eighteen years of argument later, the answer looks different.
Key Takeaway
The 1997 paper found "that the wage elasticity of daily hours of work for New York City taxi drivers is negative" and concluded their behaviour was "consistent with reference dependence"[1]. The 2015 replication, using the complete record of all trips taken in NYC taxi cabs from 2009 to 2013, reports the opposite: "drivers tend to respond positively to unanticipated as well as anticipated increases in earnings opportunities."[2]
The Verdict, Stated First
Five claims, in descending order of confidence.
One. The original finding did not survive better data. A 2015 study using every trip in New York City over five years reports drivers responding positively to higher earnings opportunities, which is the opposite sign from the 1997 result.
Two. The reason is specific and quantified. Reference dependence operates only on unanticipated variation, and the 2015 paper reports that only about one eighth of daily wage variation is unanticipated, so it can "at best, play a limited role."
Three. The intermediate rounds are the most instructive part. A 2011 paper reconciled the two positions using the challenger's own data by adding a second target, which is what a productive dispute looks like and is rarer than it should be.
Four. Our own arithmetic explains why the one-eighth figure is decisive. Even a maximal reference-dependent response on one eighth of the variation is outweighed by a modest standard response on the other seven eighths, and the observed elasticity comes out positive in three of four illustrative cases.
Five. The commercially useful conclusion is not the one the original finding suggested. If people worked less on good days, a firm raising effective hourly rates would get fewer hours. The best available evidence says the opposite, which is the ordinary economic prediction and is more reassuring than the behavioural version.
Our Grades For These Claims
Applying the scheme from the first article in this series.
Grade A for the shape of the dispute. Four papers, all in top-five economics journals, with citations and abstracts confirmed across multiple independent sources.
Grade A for the 2015 result as stated, from an abstract obtained verbatim from two independent sources including the author's own institutional repository.
Grade A for the one-eighth figure, from the author's own working paper abstract on his university's repository.
Grade B for the 1997 finding as characterised, because we did not obtain that paper and report it entirely through the 2015 paper's description of it.
Grade C for anything about magnitudes, since we obtained no elasticity estimates from any of the four papers, only their signs and directions.
Our position: this is what a scientific dispute should look like and it is the strongest evidence in this series that the process works, which is a different thing from the original finding being right.
A Note On Method
Everything here is verified to August 2026.
We obtained the 2015 paper's published abstract verbatim from an economics database and a fuller working paper abstract from the author's own university repository[2][3]. We did not obtain the paper.
We obtained the 2011 reconciliation's abstract verbatim[4] and not the paper.
We did not obtain the 1997 paper, nor the 2005 or 2008 papers, and report all three through the descriptions given by the later papers and by citation records. For a dispute this article treats as exemplary, that is a real limitation.
We obtained no effect sizes, no elasticity estimates and no confidence intervals from any paper here. Every quantitative statement below is either a reported sample characteristic, a reported fraction, or our own arithmetic.
The dataset size figures come from a university magazine article about the 2015 study, which is not an academic source and is flagged at every use[5].
All arithmetic is ours and uses invented parameters except where a published figure is named.
This article discusses research on labour supply. It is not employment, compensation or scheduling advice.
Why Cab Drivers
The reason this population became the test case, and it is a good one. Ours.
Four observations.
Most workers cannot choose their hours, which makes it impossible to observe how they would respond to a better day. Cab drivers historically could, which is the property the research needed.
A university article describes the reasoning: the author "thought that taxi drivers, who have a greater ability than most workers to set their own hours, would provide a good laboratory for examining how workers respond to earnings incentives."[5] This is a university magazine, not an academic source, flagged here and at every use.
The daily wage also varies for reasons outside the driver's control, being weather, events, and traffic, which supplies the variation a study needs without the driver choosing it.
And every trip is recorded, which is what eventually made the 2015 study possible and which almost no other occupation offers.
The 1997 Finding
Round one.
Camerer, C., Babcock, L., Loewenstein, G., and Thaler, R. (1997), Labor Supply of New York City Cabdrivers: One Day at a Time, Quarterly Journal of Economics, 112(2), 407–441[1][7].
The 2015 paper describes it: the authors "find that the wage elasticity of daily hours of work for New York City taxi drivers is negative and conclude that their labor supply behavior is consistent with reference dependence."[1] The working paper version renders the conclusion as "consistent with target earning (having reference dependent preferences)."[3]
We did not obtain the paper and report it through these descriptions.
Four observations, ours.
The author list is four of the most cited names in behavioural economics, and the paper is called "seminal" by the man who spent a decade arguing with it[1][3], which is the word an opponent uses when importance is not in dispute.
The finding is counterintuitive in exactly the way that travels. Ordinary economics predicts people work more when the hourly return is higher; this said the reverse.
The title is the argument. One day at a time means drivers treat each day as a separate accounting period rather than optimising across days, which is a mental accounting claim as much as a labour supply one.
And the mechanism named is target earning: a driver aims at a daily figure, reaches it faster on a good day, and goes home.
What A Negative Elasticity Means
Unpacking the technical term, because everything turns on its sign. Ours.
Four observations.
The wage elasticity of hours is the percentage change in hours worked for a one percent change in the effective hourly rate. Positive means more pay, more hours. Negative means more pay, fewer hours.
Standard theory expects positive, at least for daily variation, because a better day makes working the next hour more valuable while the alternative uses of that hour have not changed.
A negative elasticity is therefore an anomaly requiring explanation, and target earning supplies one: if the objective is a fixed daily sum, a higher rate reaches it sooner.
And it is costly to the worker, which is what makes it a mistake rather than a preference. Note also that the 2011 paper puts "wage" in quotation marks when describing this elasticity[4], because a cab driver has no wage. The quantity is realised earnings per hour, which the driver partly influences, and that is a subtlety the popular retelling loses.
The Arithmetic Of A Target
Making the prediction precise, because a strict target has an exact implication. Our own arithmetic, invented figures.
If a driver stops on reaching a fixed daily income, hours equal target divided by rate, and the wage elasticity of hours is exactly negative one.
With a target of $250: at $20 an hour, 12.50 hours. At $25, 10.00. At $30, 8.33. At $35, 7.14. At $40, 6.25. At $50, 5.00.
Three observations.
A twenty percent higher rate, from $25 to $30, means 16.7 percent fewer hours under a strict target.
That is a large and testable prediction, which is why the question was answerable at all. The two theories predict opposite signs, so data can distinguish them.
And a strict target is the extreme case. The empirical claim was never that the elasticity is exactly negative one, only that it is negative, which is weaker and more defensible.
Round Two: The Challenge
The response, in two papers.
Farber, H. S. (2005), Is Tomorrow Another Day? The Labor Supply of New York City Cabdrivers, Journal of Political Economy, 113(1), 46–82, and Farber, H. S. (2008), Reference-Dependent Preferences and Labor Supply: The Case of New York City Taxi Drivers, American Economic Review, 98(3), 1069–1082, June[9][10].
The 2011 paper describes the key result: "his finding that stopping probabilities are significantly related to hours but not income."[4] It also records a second line of attack, namely "Farber's criticism that estimates of drivers' income targets are too unstable to yield a useful model of labor supply."[4]
A repository copy carrying a 2005 abstract, which we report cautiously because we did not obtain the paper, describes a model where preferences depend on a reference daily income level and states "I find that there may be a reference level of income"[2], with our source truncating there.
Four observations, ours.
The stopping finding is precisely aimed. If drivers stop on hitting an income target, stopping should track income. It tracked hours instead.
Which suggests a much more ordinary explanation: drivers stop when they are tired. That is not a bias; it is a cost of effort rising through the day.
The instability criticism is separate and arguably more damaging. A target re-estimated for every driver on every day is not a theory, it is a curve fit, and that is a charge about the model rather than the data.
And note that the challenge is not a failure to replicate the elasticity. The 2011 paper attributes the negative elasticity to Camerer and colleagues, and the truncated 2005 abstract we obtained says a reference level of income may exist. By this point the elasticity was largely agreed. Its cause was not.
Round Three: The Reconciliation
The round we think is the most admirable, and the one least likely to be summarised anywhere.
Crawford, V. P., and Meng, J. (2011), New York City Cab Drivers' Labor Supply Revisited: Reference-Dependent Preferences with Rational-Expectations Targets for Hours and Income, American Economic Review, 101(5), 1912–1932, August[4][10].
Its abstract: "This paper proposes a model of cab drivers' labor supply, building on Henry S. Farber's (2005, 2008) empirical analyses and Botond Koszegi and Matthew Rabin's (2006; henceforth 'KR') theory of reference-dependent preferences. Following KR, our model has targets for hours as well as income, determined by proxied rational expectations. Our model, estimated with Farber's data, reconciles his finding that stopping probabilities are significantly related to hours but not income with Colin Camerer et al.'s (1997) negative 'wage' elasticity of hours; and avoids Farber's criticism that estimates of drivers' income targets are too unstable to yield a useful model of labor supply."[4]
Three observations, ours.
The word is reconcile, not refute. The contribution is showing that two apparently contradictory findings can both be true under a richer model.
It is estimated on the challenger's own data, which removes the most common escape route in a dispute, being that the datasets differ.
And it answers both of the objections: the stopping result and the instability criticism. A reconciliation addressing only the convenient half would be worth much less.
Targets For Hours As Well As Income
The mechanism, which is simple once stated.
The paper's own conclusion: "Our analysis builds on Farber's (2005, 2008) empirical analyses, which allowed income- but not hours-targeting and treated the targets as latent variables. Our model, estimated with Farber's data, suggests that reference dependence is an important part of the labor-supply story in his dataset, and that using KR's model to take it into account does yield a useful model of cab drivers' labor supply."[5]
Four observations, ours.
The 1997 and 2005 models have one target: income. That predicts stopping tracks income, and it did not.
The 2011 model has two, income and hours. With two targets, stopping can track hours while the elasticity stays negative, because both findings follow from one model.
The methodological move is the important one. Farber treated targets as latent variables, meaning quantities estimated from the same data they explain, which is where the instability came from. Crawford and Meng instead treat them as rational expectations and "operationalize them via sample proxies."[6]
And the proxy is specific and disciplined: each driver's "sample averages up to but not including the day in question", matched driver by driver and day-of-week by day-of-week[6]. A target built only from a driver's own past cannot be fitted to the day it is explaining.
Using The Challenger's Own Data
A methodological point worth isolating, because this series has criticised its absence elsewhere. Ours.
Three observations.
Re-estimating a rival model on the opposing side's data, while following his econometric strategies closely[6], is the strongest form a reconciliation can take. It eliminates sample, period and measurement differences in one move.
The fifty-third article covered a dispute where the parties argued about what the theory had claimed rather than about data, and the fifty-sixth covered one where the same team published both the failed replications and the meta-analysis. This is a third pattern and the cleanest of the three.
And it is only possible where data is shared, which is a fact about the field's norms rather than about anyone's insight. The replication files for the 2011 paper are themselves publicly archived with their own DOI[8].
Round Four: The Return
Eighteen years after the original.
Farber, H. S. (2015), Why you Can't Find a Taxi in the Rain and Other Labor Supply Lessons from Cab Drivers, Quarterly Journal of Economics, 130(4), 1975–2026, November, DOI 10.1093/qje/qjv026, advance access 13 July 2015[1][2].
Three observations, ours.
The paper is published in the same journal as the original, eighteen years later, which is a deliberate and appropriate choice.
It describes itself as a replication and extension and calls the target work seminal[1]. That is how this should be done.
And it began as a named lecture: the working paper records it as based on the author's Albert Rees Lecture at the annual meeting of the Society of Labor Economists, 2 May 2014[11]. A field invites someone to give an address and he uses it to overturn one of its better-known results.
Every Trip For Five Years
The data, and it is the whole story.
The abstract states it: "the complete record of all trips taken in NYC taxi cabs from 2009 to 2013."[1] The working paper: "data from all trips taken in all taxi cabs in NYC for the five years from 2009-2013."[3]
The full text describes the source: records are supplied to the city's licensing commission on a regular basis, the author "obtained full information for all trips taken in NYC taxi cabs for the five years from 2009-2013", the dataset is called TPEP, and it identifies "drivers by encrypted hack license number and medallions (cabs) by encrypted" identifier[11].
Four observations, ours.
The complete record is a phrase almost nothing in this series has been able to use. Not a sample, not a survey, not a panel. Every trip.
We did not obtain the 1997 paper and therefore cannot state its sample size, so we will not compute a ratio. What we can say is that trip sheets collected in the 1990s and a complete administrative record of five years are not the same kind of evidence.
The data is administrative rather than collected for research, which removes participation effects and recall error, and which is why it exists at all.
And this is the fifth article in this series where a finding was revisited with a dataset the original authors could not have had. In four of those five the finding changed.
The Result
The finding, in one sentence.
"In contrast, my analysis of the complete record of all trips taken in NYC taxi cabs from 2009 to 2013 shows that drivers tend to respond positively to unanticipated as well as anticipated increases in earnings opportunities."[2]
Four observations, ours.
Positively. The sign is reversed from the original. Better earnings opportunities produce more hours, not fewer, which is the standard economic prediction.
The phrase "unanticipated as well as anticipated" is the technical heart of it, and it forecloses the obvious defence. Reference dependence predicts the negative response specifically for unanticipated good days, and the paper reports a positive response there too.
Note the hedge that is genuinely present: "tend to respond." This is a statement about a central tendency across tens of thousands of drivers, not a claim that every driver behaves this way.
And we obtained no elasticity estimate, only the reported sign. We would like to tell you how positive and cannot.
One Eighth
The figure that makes the argument decisive rather than merely contrary, from the working paper abstract.
"Using the model of expectations-based reference points of Koszegi and Rabin (2006), I distinguish between anticipated and unanticipated daily wage variation and present evidence that only a small fraction of wage variation (about 1/8) is unanticipated so that reference dependence (which is relevant only in response to unanticipated variation) can, at best, play a limited role in determining labor supply."[3]
Four observations, ours.
The logic is a two-step argument and both steps are needed. First, reference dependence applies only to unanticipated variation. Second, most variation is anticipated. Therefore its scope is small.
The first step is not the author's invention. It comes from the expectations-based reference point model he names, which is the same framework the 2011 reconciliation used. He is applying the behavioural side's own model.
About one eighth is a specific quantity and it is the load-bearing number in the whole dispute. We did not obtain how it was estimated.
And "at best, play a limited role" is careful phrasing. It concedes a role and bounds it, rather than denying the phenomenon.
What That Does To The Argument
Why one eighth settles it, stated plainly before we compute it. Ours.
Three observations.
A mechanism that operates on one eighth of the variation must be very strong on that eighth to move the overall result, and it must overcome whatever the standard mechanism does on the remaining seven eighths.
The two mechanisms push in opposite directions, so they do not add. Reference dependence pushes hours down when the day is unexpectedly good; standard substitution pushes them up.
And that makes the question arithmetic rather than empirical. Given a share and two elasticities, the observed elasticity follows, which is what the next section computes.
Our Own Decomposition
Our own arithmetic, using the reported one-eighth share and invented elasticities for each component. This shows the shape of the constraint and is not a finding.
Suppose reference dependence produces its maximum effect, an elasticity of negative one, on the one eighth it can touch, while standard substitution produces a modest positive elasticity on the remaining seven eighths.
With standard elasticity +0.3: observed elasticity +0.138.
With standard elasticity +0.5: observed +0.313.
With reference dependence at only negative 0.5 and standard at +0.3: observed +0.200.
And with standard elasticity at zero, meaning no substitution response at all: observed negative 0.125.
Four observations.
Three of the four cases produce a positive observed elasticity, which is the sign the 2015 paper reports.
The only negative case requires the standard response to be exactly zero, which nobody in this dispute claims.
So the one-eighth figure does not merely weaken reference dependence. It makes the originally observed negative elasticity difficult to generate at all unless the standard channel is switched off.
And every input here except the one eighth is ours and invented. The exercise shows why the share matters, not what the true elasticities are.
What Counts As Anticipated
The strongest objection to the strongest evidence in this article, which we raise because none of our sources does and because the whole 2015 argument rests on it. Ours.
The one-eighth figure is not a measurement. It is the output of a classification: daily wage variation is split into a part drivers could have anticipated and a part they could not, and the split is produced by a model the researcher chose.
Four observations.
The model is named in the abstract, being expectations-based reference points[3], so the classification is disciplined rather than arbitrary. But it is still a model, and a different one could apportion the variation differently.
The direction of any error is knowable in advance. If the model treats as anticipated something drivers actually did not anticipate, the unanticipated share is understated, and reference dependence is given less room than it deserves. If it errs the other way, the share is overstated.
Our own decomposition shows how much rides on this. At one eighth the observed elasticity comes out positive in three of four illustrative cases, and raising the share changes that, so a reader who doubts the classification should doubt the conclusion in proportion.
And the honest position is that we cannot evaluate the classification because we did not obtain the paper. We flag it as the place a serious reader would look first, and we note that the headline finding does not depend on it entirely: the reported positive response to unanticipated increases is a direct result that stands whatever share those increases represent.
The Escalation
What each round had to work with, which is the pattern worth carrying away. Ours.
1997: trip sheets, hand-collected. We did not obtain the sample size.
2005 and 2008: larger trip sheet data, described by later papers as the basis for both a challenge and a reconciliation. We did not obtain the samples.
2011: no new data at all, a new model estimated on the 2005 and 2008 data.
2015: the complete administrative record of every trip over five years, reported elsewhere as more than 700 million trips and some 62,000 drivers.
Three observations.
The dispute was settled by data availability rather than by argument, and the data became available because taxis were required to record every trip electronically.
Round three is the exception and it is instructive. A better model on the same data produced genuine progress, which is worth noting for anyone who thinks only more data helps.
And the honest reading of the sequence is not that the original authors were careless. They used the best data that existed in 1997, and the finding stood for eight years before anyone could challenge it properly.
And A Randomised Experiment
A separate line of evidence, from a different occupation and a stronger design.
Reference lists identify Fehr, E., and Goette, L. (2007), Do Workers Work More if Wages Are High? Evidence from a Randomized Field Experiment, American Economic Review, 97(1), 298–317, March[9].
We did not obtain this paper and report its title and citation only.
Three observations, ours.
A randomised field experiment on this question is a stronger design than any observational study, because the wage variation is assigned rather than occurring naturally.
The title poses the question in its simplest form, and we cannot tell you the answer, which is unsatisfying for the section that would otherwise be the strongest evidence here.
A teaching problem set we obtained frames it as contested: it asks students to evaluate the statement that this study's finding "contradicts the reference dependence story" and instructs them to "discuss"[10]. That is a course exercise rather than a source, and we report it only as evidence that the interpretation is not straightforward.
Divers, Vendors And Fishermen
Where else the question has been asked, from citation records.
Reference lists identify Oettinger, G. S. (1999), An Empirical Analysis of the Daily Labor Supply of Stadium Vendors, Journal of Political Economy, 107(2), 360–392; Eggert, H., and Kahui, V. (2013), Reference-dependent behaviour of paua (abalone) divers in New Zealand, Applied Economics, 45(12), 1571–1582; and Hammarlund, C. (2018), A trip to reach the target? The labor supply of Swedish Baltic codfishermen, Journal of Behavioral and Experimental Economics, 75, 1–11[11][12].
We obtained none of these papers and report titles and citations only.
Three observations, ours.
The occupations share the property that made cab drivers useful: self-set hours and externally varying returns. Stadium vendors, abalone divers and cod fishermen all choose when to stop.
The geographic spread is genuine, covering the United States, New Zealand and Sweden, which is more than most literatures in this series can show.
And we grade none of it, because a list of titles is not evidence and two of the three titles are phrased as questions rather than findings.
What Actually Survives
Our reading, stated directly.
Five statements.
The best available evidence says people work more when the return is higher, which is the ordinary prediction and the opposite of the famous finding.
The negative elasticity was real in the earlier data and was found by the challenger too. What changed was the data and the decomposition, not anyone's competence.
Reference dependence is bounded rather than eliminated. The 2015 paper concedes it a limited role, on the grounds that it can only act on about one eighth of the variation.
The 2011 reconciliation remains standing. Nothing in the 2015 paper's abstract addresses whether a two-target model fits the older data, and it was estimated on data the 2015 paper does not use.
And we obtained no magnitudes anywhere, so every claim here is about signs and shares.
Your Own Hours
The first application, and the most personal for an owner-operator. Ours, untested, and not employment advice.
Four points.
The original finding, if true, would describe a genuine and expensive error: working long hours on unproductive days and short hours on productive ones. Over a year that costs real money for no additional leisure.
The best evidence says people do not systematically do this, which is reassuring rather than actionable.
But the diagnostic is cheap and worth running anyway. Correlate your own hours against your own daily or weekly revenue for a year. A negative relationship is the pattern the 1997 paper described, and it is your data rather than anyone's theory.
And the target-setting practice is the thing to watch. A daily or monthly revenue target is exactly the mechanism the original paper proposed, and the fifty-seventh article's finding on surrogation applies: a target set as a proxy can become the thing being pursued.
Commission And Piece Rates
The second application, where the sign of the elasticity is a design question. Ours.
Four points.
Any commission or piece-rate scheme is a bet that a higher effective hourly return produces more effort. That is a positive elasticity assumption, and it is the one the 2015 evidence supports.
If the 1997 finding had held generally, raising commission rates would have reduced hours worked, which would make the entire structure counterproductive. The evidence does not support that worry.
The residual concern is narrower and worth keeping. An explicit target inside a commission scheme reintroduces the mechanism, because a person paid on a target rather than a rate has a reason to stop on reaching it.
And that suggests a design preference we would state cautiously. Rates are less likely to produce stopping behaviour than thresholds, which follows from the theory rather than from any evidence we obtained.
Your Contractors
The third application, and the one closest to the studied population. Ours.
Three points.
Self-employed contractors are the group in your supply chain with genuinely self-set hours, which is the property that made cab drivers the test case.
If you pay a contractor more per hour or per unit, the evidence suggests you get more of their time, not less. That is the useful direction for anyone competing for a scarce trade.
And the counterpart is worth naming. A contractor with a fixed weekly income requirement behaves like the target earner of the original model, and the way to find out is to raise the rate and watch whether availability rises or falls.
And The Rain
The title's question, which we can answer only partly and would rather flag than fudge.
The 2015 paper's title asks why you cannot find a taxi in the rain, and a university article describes the study as combining trip data "with information about the weather"[5].
Three observations, ours.
The original explanation, implied by the 1997 finding, is that rain raises the hourly rate so drivers hit their target and go home, leaving fewer cabs on the road.
The 2015 result rules that out, since drivers respond positively to better earnings opportunities. So the target-earning explanation of the rain phenomenon does not survive its own author's data.
And we did not obtain the paper's actual explanation. The title promises one, the abstract does not state it, and this is the most conspicuous gap in this article. A reader who wants the answer needs the paper.
The Lesson Is About Evidence
What we would take from this beyond the finding itself. Ours.
Four observations.
Every round was conducted in public, in top journals, by named people using each other's data. Nobody wrote a commentary; they wrote papers.
The dispute took eighteen years and required a technology change to settle. That is not a failure of the participants and it is a reason to hold recent findings loosely.
The reconciliation round produced genuine progress with no new data at all, which is the counterexample to the assumption that more data is the only way forward.
And the original authors' finding was correct about the data they had. The fiftieth article in this series found that the most common structure it encountered was a famous result reproducible from something other than its stated mechanism, and this is a cleaner case: a real pattern in a small sample that a larger sample reversed.
What To Do
Do not assume people work less when the return is higher. The best available evidence, from the complete record of five years of New York taxi trips, reports the opposite.
Check your own hours against your own revenue. A negative relationship over a year is the pattern the original paper described, and it is your data rather than anyone's theory.
Prefer rates to thresholds in incentive design. A target gives a reason to stop on reaching it; a rate does not. That follows from the theory rather than from evidence we obtained.
Notice when a mechanism can only touch part of the variation. Reference dependence acts only on unanticipated changes, and about one eighth of daily wage variation is unanticipated, which bounds what it can explain.
Ask what data a finding was based on and what data exists now. This finding stood for eight years, was reconciled at fourteen, and was reversed at eighteen when the complete record became available.
Value reconciliations, not just refutations. The 2011 paper resolved an apparent contradiction using the opposing side's data and a richer model, which is the most useful thing anyone did in this sequence.
Hold the 2015 result loosely too. It is one paper, we obtained only its abstract, its central figure rests on a researcher-chosen classification, and the dispute has reversed twice already.
Do not read this as a general verdict on behavioural economics. A bounded role is not no role, and the paper's own language concedes one.
The Limits Of This Analysis
Several caveats matter, and the sourcing here is thinner than the article's confidence might suggest. This article discusses research on labour supply and is not employment, compensation or scheduling advice; the applications are our own reasoning and untested. Everything is verified to August 2026. We did not obtain a single one of the four principal papers. We have the 2015 paper's published abstract and a fuller working paper abstract from the author's own university repository; the 2011 paper's abstract; and nothing at all of the 1997, 2005 or 2008 papers beyond how later papers describe them. For a dispute this article calls exemplary, that is a real limitation and we would rather state it prominently than in a footnote. We obtained no elasticity estimates, no effect sizes and no confidence intervals from any paper, so every quantitative claim here is a reported share, a reported sample characteristic, or our own arithmetic. We did not obtain how the one-eighth figure was estimated, and it is the load-bearing number in the whole dispute; as the body sets out, it is a classification produced by a chosen model rather than a direct measurement, and we could not evaluate it. The dataset size figures of 700 million trips and 62,000 drivers come from a university magazine article, not an academic source, flagged at every use. We did not obtain the randomised field experiment, which would be the strongest evidence here, and can report only its title; a teaching problem set we obtained treats its interpretation as contested. We obtained none of the studies in other occupations and report titles only. All arithmetic is ours. The target-earning table uses an invented income target, and the elasticity decomposition uses the reported one-eighth share with entirely invented component elasticities; it demonstrates why the share constrains the argument and is not an estimate of anything. And we could not obtain the 2015 paper's actual explanation of the rain phenomenon its title promises, which is the most conspicuous gap in this article.
Frequently Asked Questions
What was the original finding?
Did it hold up?
Why does the one-eighth figure matter?
Is that one-eighth figure solid?
What did the 2011 paper do?
Why can't you find a taxi in the rain?
What should a business owner take from this?
References
- Publisher record for the 2015 replication, reproducing the opening of its abstract on the seminal work of Camerer and colleagues finding that the wage elasticity of daily hours of work for New York City taxi drivers is negative, and concluding that their labor supply behavior is consistent with reference dependence; and confirming the 2015 citation as Henry S. Farber, Why you Can't Find a Taxi in the Rain and Other Labor Supply Lessons from Cab Drivers, The Quarterly Journal of Economics, Volume 130, Issue 4, November 2015, Pages 1975–2026, DOI 10.1093/qje/qjv026. Note: the publisher's record. Our source for the characterisation of the 1997 finding, which we report through the later paper's description because we did not obtain the 1997 paper. academic.oup.com
- Economics database record for Farber, H. S. (2015), Why you Can't Find a Taxi in the Rain and Other Labor Supply Lessons from Cab Drivers, The Quarterly Journal of Economics, 130(4), 1975–2026, reproducing the abstract in full: on the author replicating and extending the seminal work of Camerer and colleagues, who find that the wage elasticity of daily hours of work for New York City taxi drivers is negative and conclude their labor supply behavior is consistent with reference dependence; and on the author's analysis of the complete record of all trips taken in NYC taxi cabs from 2009 to 2013 showing, in contrast, that drivers tend to respond positively to unanticipated as well as anticipated increases in earnings opportunities. Note: an economics database record reproducing the published abstract. We obtained the abstract only and not the paper, and report no elasticity estimate. ideas.repec.org
- Working paper version of the same study, hosted on the author's own university economics repository, reproducing a fuller abstract: on Camerer, Babcock, Loewenstein and Thaler (1997) finding in a seminal paper that the wage elasticity of daily hours of work of New York City taxi drivers is negative and concluding their labor supply behavior is consistent with target earning, having reference dependent preferences; on the author replicating and extending that analysis using data from all trips taken in all taxi cabs in NYC for the five years from 2009 to 2013; on the author using the model of expectations-based reference points of Koszegi and Rabin (2006) to distinguish between anticipated and unanticipated daily wage variation; and on the author presenting evidence that only a small fraction of wage variation, about one eighth, is unanticipated, so that reference dependence, which is relevant only in response to unanticipated variation, can at best play a limited role in determining labor supply. Note: the author's own university repository. Our source for the one-eighth figure, which is the load-bearing number in this dispute; we did not obtain how it was estimated and the body sets out why that matters. repec-prod.princeton.edu
- Repository page carrying the abstract of Crawford, V. P., and Meng, J., New York City Cabdrivers' Labor Supply Revisited: Reference-Dependent Preferences with Rational-Expectations Targets for Hours and Income: on the paper reconsidering whether cabdrivers' labor supply decisions reflect reference-dependent preferences; on the authors, following Botond Koszegi and Matthew Rabin (2006), constructing a model with targets for hours as well as income, both determined by rational expectations; and on the authors, estimating using Henry S. Farber's 2005 and 2008 data, showing that the reference-dependent model can reconcile his 2005 finding that drivers' stopping probabilities are significantly related to hours but not income with the negative wage elasticity of hours found by Colin Camerer and colleagues (1997) and Farber (2005, 2008). Note: a repository page reproducing the abstract. We obtained the abstract only and not the paper; this is also our only source for the characterisation of Farber's 2005 finding. researchgate.net
- University alumni magazine article about the 2015 study, recording that the author studies employment and thought taxi drivers, who have a greater ability than most workers to set their own hours, would provide a good laboratory for examining how workers respond to earnings incentives; and that he obtained records of every cab ride taken in New York City between 2009 and 2013, a database of more than 700 million trips involving some 62,000 drivers, combining some of that data with information about the weather. Note: a university alumni magazine, not an academic source, flagged at every use. Our only source for the dataset size figures and for the study's stated motivation. paw.princeton.edu
- Economics database record confirming Colin Camerer, Linda Babcock, George Loewenstein and Richard Thaler (1997), Labor Supply of New York City Cabdrivers: One Day at a Time, The Quarterly Journal of Economics, volume 112(2), pages 407–441; and recording an earlier version as a 1996 working paper of the California Institute of Technology, Division of the Humanities and Social Sciences. Note: a database record used to confirm the citation independently. We did not obtain the paper and report its findings only through later papers' descriptions. ideas.repec.org
- Behavioural economics reference bibliography confirming Farber, H. S. (2005), Is tomorrow another day? The labor supply of New York City cabdrivers, Journal of Political Economy, 113(1), 46–82, DOI 10.1086/426040; Farber, H. S. (2008), Reference-Dependent Preferences and Labor Supply: The Case of New York City Taxi Drivers, American Economic Review, 98(3), 1069–1082, DOI 10.1257/aer.98.3.1069; Farber, H. S. (2015), The Quarterly Journal of Economics, 130(4), 1975–2026, DOI 10.1093/qje/qjv026; and Camerer, C., Babcock, L., Loewenstein, G., and Thaler, R. (1997), The Quarterly Journal of Economics, 112(2), 407–441, DOI 10.1162/003355397555244. Note: a reference bibliography, used to confirm citations and DOIs for all four principal papers independently. We obtained none of them. behaviouraleconomics.jasoncollins.blog
- Economics database record confirming Vincent P. Crawford and Juanjuan Meng (2011), New York City Cab Drivers' Labor Supply Revisited: Reference-Dependent Preferences with Rational-Expectations Targets for Hours and Income, American Economic Review, volume 101(5), pages 1912–1932, August. Note: a database record used to confirm the 2011 citation, journal, volume, pages and month independently. ideas.repec.org
- Reference list carried on a publisher's reference work entry on reference points and effort, confirming Fehr, E., and Goette, L. (2007), Do workers work more if wages are high? Evidence from a randomized field experiment, American Economic Review, 97(1), 298–317; Goette, L., Huffman, D., and Fehr, E. (2004), Loss aversion and labor supply, Journal of the European Economic Association, 2(2–3), 216–228; and Farber (2008) and Farber (2015). Note: a reference list; citations only. We did not obtain the randomised field experiment, which would be the strongest evidence in this article, and report only its title. link.springer.com
- University course problem set on reference dependence, citing Farber (2015) and Fehr and Goette (2007), and posing to students the statement that the latter find bike messengers work more shifts when paid more and that this contradicts the reference dependence story, with the instruction to discuss whether that is right. Note: a teaching problem set, not a source. Reported only as evidence that the interpretation of that experiment is treated as contested rather than settled; nothing in this article relies on it for a factual claim. econ.berkeley.edu
- Economics working paper database record for a study of reference-dependent behaviour among abalone divers, carrying a reference list confirming Oettinger, G. S. (1999), An Empirical Analysis of the Daily Labor Supply of Stadium Vendors, Journal of Political Economy, 107(2), 360–392; Farber (2005) and Farber (2008); and recording Eggert, H., and Kahui, V. (2013), Reference-dependent behaviour of paua (abalone) divers in New Zealand, Applied Economics, 45(12), 1571–1582. Note: a database record; citations only. We obtained none of the studies named. ideas.repec.org
- Economics database record for research on labour supply and reference dependence, carrying a reference list confirming Farber (2015), Quarterly Journal of Economics, 130(4), 1975–2026; Fehr and Goette (2007), American Economic Review, 97(1), 298–317; Goette, Huffman and Fehr (2004); and Hammarlund, C. (2018), A trip to reach the target? The labor supply of Swedish Baltic codfishermen, Journal of Behavioral and Experimental Economics, 75, 1–11. Note: a database record; citations only. We obtained none of the studies named. ideas.repec.org
This article discusses research on labour supply and is not employment, compensation or scheduling advice. None of the four principal papers was obtained. The 2015 paper is reported from its published abstract and a fuller working paper abstract; the 2011 paper from its abstract; and the 1997, 2005 and 2008 papers only through later papers' descriptions. No elasticity estimate, effect size or confidence interval from any paper is reported. The one-eighth figure is a classification produced by a chosen model rather than a direct measurement, and could not be evaluated. The dataset size figures come from a university magazine article, not an academic source. All arithmetic is the authors' own and uses invented parameters except the reported one-eighth share.