The eighty-ninth article was about a claim whose evidence we could not obtain. This one is about a claim whose evidence a team of researchers went looking for and reported was not there, and about the specific shape of the thing they could not find.

Key Takeaway

The 2008 review concludes that "very few studies have even used an experimental methodology capable of testing the validity of learning styles applied to education," and that of those which did, "several found results that flatly contradict the popular meshing hypothesis."[2] On our own arithmetic, three partial designs give the same answer whether the hypothesis is true or false, and one multiplication on a four-cell table settles it.

The Verdict, Stated First

Five claims, in descending order of confidence.

One. The review's conclusion is unambiguous and we quote it in full. There is, on its finding, no adequate evidence base to justify incorporating learning-styles assessments into general educational practice.

Two. The failure is one of study design rather than of results. The core evidence is described as missing, not as negative, which is a different and more interesting problem.

Three. The authors added a caveat that is almost never quoted, and we quote it: it would be an error to conclude that all possible versions have been tested and found wanting.

Four. The required test has an exact form, and on our own arithmetic it reduces to multiplying two numbers and checking the sign.

Five. Acting on the unvalidated version can make training worse rather than merely wasteful, because on our own arithmetic it routes half your people to the method that is worse for everybody.

Our Grades For These Claims

Applying the scheme from the first article in this series, and the sourcing here is much stronger than the previous article's.

Grade A for the review's definition and hypothesis statement, from an abstract obtained verbatim from a national medical library database.

Grade A for its conclusion and caveat, obtained verbatim from a university repository record.

Grade A for the paper's own section headings, obtained from a copy hosted on an academic networking site, which tell you the structure of its argument.

Grade C for the three stated conditions of a valid test, which reach us through an advocacy organisation's summary.

Grade D for the subsequent literature, which reaches us as citations in a reference list on a non-academic page.

Grade A for our own arithmetic, which is exact, and which contains one section we withdrew as circular, documented below.

A Note On Method

Everything here is verified to August 2026.

We obtained the abstract verbatim from a national medical library database[1] and the conclusion and its caveat verbatim from a university repository[2].

We obtained the paper's table of contents and passages of its body from a copy hosted on an academic networking site[3].

Two sources are non-academic: an education advocacy organisation[4] and a professional council's article[5], both flagged at every use.

We did not obtain the full review, so we cannot say how many studies it examined or name the ones it found used appropriate methods.

We obtained none of the subsequent literature, which reaches us as citations only.

All arithmetic is ours and every figure in it is invented. This article discusses research on instruction and is not training or educational advice.

The Definition

What is being claimed, in the review's own words.

Its abstract records: "The term 'learning styles' refers to the concept that individuals differ in regard to what mode of instruction or study is most effective for them. Proponents of learning-style assessment contend that optimal instruction requires diagnosing individuals' learning style and tailoring instruction accordingly."[1]

And on how the assessments work: "Assessments of learning style typically ask people to evaluate what sort of information presentation they prefer (e.g., words versus pictures versus speech) and/or what kind of mental activity they find most engaging or congenial (e.g., analysis versus listening), although assessment instruments are extremely diverse."[1]

Four observations, ours.

The claim has two parts, and separating them is most of the work. One is that people differ in what suits them. The other is that you should diagnose this and act on it.

The first part is nearly uncontroversial and is not what the evidence problem concerns. People obviously differ.

The word "prefer" in the second quotation is the crux. The instruments measure what people say they like, and liking a format is a different thing from learning better from it.

And "extremely diverse" is a real complication the review flags in its own abstract. There is no single instrument, so a study validating one says little about another.

Two consequences of that diversity, ours. A firm buying an assessment is buying one instrument among many, and evidence about any other one does not transfer to it.

And it makes the literature harder to pool than it looks. Studies using different instruments are not testing the same construct, which is a difficulty the review presumably had to work around and which we did not obtain its treatment of.

The Meshing Hypothesis

The specific testable claim, which the review names and defines.

Its abstract records: "The most common, but not the only, hypothesis about the instructional relevance of learning styles is the meshing hypothesis, according to which instruction is best provided in a format that matches the preferences of the learner (e.g., for a 'visual learner,' emphasizing visual presentation of information)."[1]

Four observations, ours.

Naming the hypothesis is what makes the review useful rather than merely sceptical. A vague idea cannot be tested; a named one can, and everything that follows depends on this definition.

The phrase "but not the only" is a qualification the authors put in their own abstract, and it recurs in their conclusion. They are careful throughout not to claim more than they establish.

The hypothesis is comparative and directional. It does not say visual learners benefit from visual instruction; it says they benefit more than other learners do, which is a much stronger and more specific claim.

And that specificity is what determines the study design required, which is the substance of this article.

One further consequence of the comparative form, ours. The hypothesis can be false while every part of the folk version is true. People can differ, have real preferences, and learn better from good visuals, and matching can still be worthless.

That is why the argument is hard to have. Each component of the belief survives inspection and the conclusion does not follow from them, which is a considerably more awkward position than a simple falsehood.

The 2008 Review

The source.

Pashler, H., McDaniel, M., Rohrer, D., and Bjork, R. (2008), Learning Styles: Concepts and Evidence, Psychological Science in the Public Interest, 9(3), 105–119, December, DOI 10.1111/j.1539-6053.2009.01038.x[1].

A non-academic source records that the review was commissioned by a professional psychological association[5], which we could not verify independently.

Four observations, ours.

The journal's name is unusual and matters. Psychological Science in the Public Interest exists specifically to assess whether practices with public consequences have evidence behind them, which is a different remit from a normal research journal.

Its articles are commissioned reviews rather than submitted findings, on that description, which means the authors were asked the question rather than choosing to argue a position.

The four authors are cognitive psychologists working on learning and memory, so this is the field the claim belongs to assessing its own applied practice.

And the paper's own section headings, which we obtained, tell you the structure before you read a word of it.

Two of those headings are worth naming here because they frame everything below. "Interactions as the Key Test of the Learning-Styles Hypothesis" announces that the paper is about study design rather than about results[3].

And "Costs and Benefits of Educational Interventions"[3] announces that the authors intend to treat this as a resource-allocation question, which is the frame a business reader will find most usable.

The Conclusion

What they found, quoted at length because the wording is precise.

The review states: "Although the literature on learning styles is enormous, very few studies have even used an experimental methodology capable of testing the validity of learning styles applied to education. Moreover, of those that did use an appropriate method, several found results that flatly contradict the popular meshing hypothesis."[2]

And: "We conclude therefore, that at present, there is no adequate evidence base to justify incorporating learning-styles assessments into general educational practice. Thus, limited education resources would better be devoted to adopting other educational practices that have a strong evidence base, of which there are an increasing number."[2]

Four observations, ours.

"Enormous" and "very few" in the same sentence is the finding in miniature. A vast literature that mostly could not answer the question it was about.

The review's own section heading for this is "Style-by-Treatment Interactions: The Core Evidence Is Missing"[3], which is as direct as a heading gets.

A second heading reads "Learning-Styles Studies With Appropriate Methods and Negative Results"[3], so the studies that were properly designed exist and are a small set.

And the final sentence is the one a business should read twice. The recommendation is not to do nothing; it is to spend the same budget on practices with an evidence base, of which the review says there is an increasing number.

One point about the strength of the wording, ours. "Flatly contradict" is unusually blunt for a review, and it applies specifically to the small set of studies that were designed properly.

Which produces an uncomfortable pattern for the hypothesis rather than a neutral one. The evidence is mostly absent, and where it is present it mostly points the wrong way, on the review's own account.

The Caveat They Added

The sentence that almost never survives into the popular version of this debunking, and which we quote for the same reason we quoted the subtle-effects sentence in the eighty-sixth article.

The review states: "However, given the lack of methodologically sound studies of learning styles, it would be an error to conclude that all possible versions of learning styles have been tested and found wanting; many have simply not been tested."[2]

Four observations, ours.

Absence of evidence is not evidence of absence, and the authors say so themselves, in their own conclusion, about their own finding.

That distinction is doing real work here. The review reports that the studies were not run, not that they were run and failed, and those support different strengths of conclusion.

Which means anyone citing this review as proof that learning styles do not exist has committed the same category of error the review documents, in the opposite direction.

And it leaves the practical position clear even so. An untested claim is not a basis for spending money, which is what the recommendation says, and is a weaker and more defensible statement than declaring the claim false.

What Evidence Would Be Required

The constructive part, and the reason this review is worth more than a debunking.

An advocacy organisation summarising the review records that experiments testing the hypothesis "would have to satisfy three conditions," beginning with learners being allocated to two or more groups on the basis of an assessment of their presumed learning style[4]. This is a non-academic source, flagged.

A second non-academic source states the requirement more precisely: that any credible validation "must demonstrate a particular type of statistical result known as a 'crossover interaction'"[5]. Also non-academic, flagged.

The paper's own body confirms the term, referring to "the style-by-method crossover interaction described previously"[3].

Four observations, ours.

The review does not merely say the evidence is missing. It specifies exactly what the missing evidence would look like, which makes the claim falsifiable and the criticism fair.

A crossover interaction is a specific and visualisable pattern, not a statistical abstraction, and the next several sections show it.

It is also a test a business reader can apply, to this claim and to others of the same shape, which is why we have built the article around it.

And the paper's own phrasing distinguishes the crossover requirement from the meshing hypothesis proper, noting a weaker version "requires only the style-by-method crossover interaction" without requiring that the optimal method match each group's stated style[3]. That is a subtlety we obtained and did not fully resolve, and we flag it.

The Four Cells

Our own arithmetic from here, on entirely invented scores.

Testing the claim needs a two-by-two: two learner types, two teaching methods, and all four combinations measured. Visual learners taught visually, visual learners taught verbally, verbal learners taught visually, and verbal learners taught verbally.

Four observations.

All four cells are necessary, and this is the whole of the design problem. A study missing any one of them cannot produce the required result.

That is a demanding requirement in practice. It means teaching some people in the way you believe is wrong for them, which is uncomfortable and is presumably part of why it was rarely done.

It is also expensive, needing four groups where an intuitive study needs one or two, with the sample size to support comparisons between them.

And with the four cells in hand, the three scenarios below become distinguishable, which is exactly what the incomplete designs cannot do.

Scenario A: No Meshing At All

Ours, invented scores out of 100.

Suppose visual teaching is simply better for everyone. The four cells read:

Visual learners: visual method 80, verbal method 70.

Verbal learners: visual method 80, verbal method 70.

Four observations.

Every row is identical. Learner type makes no difference to anything, and the meshing hypothesis is completely false in this world.

But notice the top-left cell. Visual learners given visual instruction score 80, which is their best cell, and a study reporting that would be reporting a true fact.

They also score higher than they do with verbal instruction, by ten points. Another true fact, and it supports the hypothesis not at all, because the verbal learners show exactly the same gap.

And this scenario is entirely plausible in the real world. Diagrams and demonstrations help most people on most technical material, which would produce this pattern with no learning styles involved.

Two reasons that plausibility matters more than it might seem, ours. The false scenario is not a contrived edge case; it is what you would expect if presentation format simply has general effects, which is the default assumption.

And it means the burden of proof sits where the review puts it. A general effect explains the observations at least as well as matching does, and the crossover is the only observation that separates them.

Scenario B: A Genuine Crossover

Ours, same invented scale.

Visual learners: visual 80, verbal 70.

Verbal learners: visual 70, verbal 80.

Four observations.

The rows point in opposite directions. Visual learners do ten points better with visual instruction; verbal learners do ten points worse with it.

This is the pattern the hypothesis predicts, and only this pattern supports it. Matching the method to the learner produces a real gain, and mismatching produces a real loss.

Plotted, the two lines cross, which is where the name comes from. Neither method is better in general; each is better for its own group, and that is precisely the claim.

And this is a strong prediction, which is a virtue. It is easy to fail, and a hypothesis that is easy to fail and does not is worth something.

One thing this scenario would also imply, ours, and it is worth stating because it is rarely noticed. Under a genuine crossover, the average across everyone is identical for both methods. Seventy-five either way.

Which means a study reporting that one method beat the other overall has, in that finding alone, evidence against a pure crossover. The signature of real matching is that neither method wins in aggregate.

Scenario C: An Interaction That Is Not Enough

The case most likely to be mistaken for support. Ours, same scale.

Visual learners: visual 85, verbal 70.

Verbal learners: visual 80, verbal 70.

Four observations.

There is a real interaction here. The benefit of visual teaching is larger for visual learners, fifteen points against ten, so learner type does modify the effect.

But visual teaching is still better for both groups. Nobody should be taught verbally on this evidence, and matching would send half your people to the worse method.

This pattern is called an ordinal interaction, and it is genuinely different from a crossover, which is why the review specifies the latter.

And a study finding this would be reported, honestly, as evidence that learning style modifies instructional effectiveness. That sentence would be true and would not support matching, which is the trap this article exists to name.

Two ways this pattern gets described in practice, ours, and both are defensible English. "Learning style significantly moderated the effect of presentation format" is accurate here, and a reader will hear it as support.

So will "visual learners benefited most from visual instruction." Also accurate, also compatible with everyone benefiting from visual instruction, and the word doing the damage is "most" rather than anything false.

What A Partial Test Concludes

Why the incomplete designs mislead. Ours, using the three scenarios above.

Suppose a study measures only visual learners, teaching them both ways, and asks whether they did better with visual instruction.

In scenario A, where the hypothesis is false: yes. In scenario B, where it is true: yes. In scenario C, where it is false: yes.

Four observations.

The same answer in all three, because the design never looks at the other row.

A study of this shape is not weak evidence for the hypothesis. It is no evidence at all, since its result is identical whether the hypothesis holds or not.

And the weakest design is worse still. Teaching matched groups and measuring improvement finds improvement in every scenario, including one where learner type is irrelevant to everything.

Which is the eighty-sixth article's problem in a new setting. People improve; the question is whether they improved because of the thing you did, and a design with no comparison cannot say.

A Correction To Our Own Section

Documented rather than removed, in this series' usual habit.

We first wrote a section computing how often a partial test would falsely confirm the hypothesis, and printed a table of probabilities.

The table was circular. We had defined the false confirmation rate as the probability that the matched method happened to be the generally better one, and then printed those two quantities as separate columns showing identical numbers. That is a restatement of a definition, not a result.

Three observations, ours.

The tell was visible in the output and we nearly missed it. Two columns of a table containing the same numbers means one was computed from the other, which is the same signal that caught the identity in the eighty-second article.

We replaced it with the question the section should have asked, which is not how often a bad design misleads but which designs can distinguish the scenarios at all. That is answerable and appears below.

And the replacement is more useful than the original would have been, which has now happened often enough that we would state it as a rule. A section that has to be rebuilt usually improves, because the second attempt is aimed at a question rather than at a conclusion.

Which Designs Can Tell Them Apart

Our own arithmetic, applying four designs to the three invented scenarios.

Matched groups only, measured before and after. Result in A: improves. In B: improves. In C: improves. Cannot distinguish anything.

Visual learners taught both ways. In A: visual wins. In B: visual wins. In C: visual wins. Cannot distinguish anything.

Both learner types, taught visually. In A: a tie. In B: visual types score higher. In C: visual types score higher. Separates A from the others, and cannot separate B from C, which is the distinction that matters.

The full two-by-two. In A: no crossover. In B: crossover. In C: no crossover. Separates all three.

Four observations.

Only the complete design answers the question, and the three partial designs are not weaker versions of it but different questions with different answers.

The third design is the interesting near-miss. It distinguishes the completely null case and still cannot tell you whether to match, because B and C look identical to it.

Which explains how an enormous literature could accumulate without settling anything. Each study answered a real question, and none of them answered this one.

And it is worth saying that nothing here implies bad faith. The partial designs are cheaper, more humane and more intuitive, and the reason to run the awkward one is only that it is the one that works.

One further constraint worth naming, ours. The complete design requires deliberately teaching people in the way you believe is worst for them, and in a school or a workplace that is a real objection rather than a squeamish one.

Which is presumably part of the answer to how an enormous literature avoided the necessary study. The design that answers the question is the one hardest to get approved, and that is a structural problem rather than anybody's failing.

The One-Line Test

The whole thing reduced to an operation you can perform on any four-cell table. Ours.

Take the difference between the two methods for the first group. Take the difference between the two methods for the second group. Multiply them. If the product is negative, you have a crossover.

In scenario A: 10 times 10 equals plus 100. No crossover.

In scenario B: 10 times minus 10 equals minus 100. Crossover.

In scenario C: 15 times 10 equals plus 150. No crossover.

Four observations.

The test works because a crossover means the two differences point in opposite directions, and a product of opposite signs is negative. That is the entire mathematical content.

It is arithmetic rather than statistics, and a reader with a table and no software can apply it in seconds.

What it does not do is tell you whether the crossover is statistically reliable, which needs sample sizes and a significance test, and which the eighty-eighth article's figures suggest is a serious obstacle at small samples.

But it settles the prior question, which is the one usually skipped. Before asking whether a difference is significant, check whether it is the right shape, because a significant ordinal interaction still does not support matching.

The Cost Of Acting On It

What matching costs if scenario A is the true one. Ours, invented figures, forty staff with half classified as verbal learners.

Everyone taught visually: average score 80.0.

Everyone matched to their style: average score 75.0.

A cost of 5.0 points, before counting anything else.

Four observations.

Matching made the training worse, not merely more expensive, by routing half the group to the method that is worse for everybody.

That is the specific danger and it is worth separating from the usual complaint. A useless intervention wastes money; this one degrades the outcome, if the world looks like scenario A.

And the other costs are real and additional. The assessment instrument, the time spent completing it, and building two versions of every training module, none of which appear in the five points.

We would flag that this depends entirely on scenario A being the true one, which we do not know and neither does anyone else, which is the review's actual finding.

Two ways to hold that uncertainty sensibly, ours. Under scenario B, matching gains you the same five points it loses under scenario A, so the stake is symmetric and the question is only which world you are in.

And under the review's account nobody can currently tell, which makes the cheaper and simpler option the reasonable default, since it costs nothing extra to be wrong about.

What Actually Survives

Our reading, stated directly.

Five statements.

The review found the core evidence missing rather than negative, which is its own section heading and is a narrower claim than a refutation.

The few studies with appropriate methods are described as several finding results that flatly contradict the meshing hypothesis, which is the strongest negative statement in the paper.

The authors explicitly decline to conclude that all versions have been tested and found wanting, and we have quoted that.

The required design is a crossover interaction, which is checkable by multiplying two numbers.

And the recommendation is to redirect the same resources to practices with an evidence base, not to abandon training.

Which is a narrower and more useful conclusion than the one that usually circulates, ours. The review is not an argument that training does not work. It is an argument about which of two things to buy with the same money.

Why It Persists

Our reading, offered as reasoning rather than evidence.

Four observations.

The first half of the claim is true and the second does not follow. People do differ, and they do have preferences, and neither establishes that matching helps.

It is also flattering and inclusive, which the review's own section on appeal presumably addresses and which we did not obtain. Telling somebody they learn differently rather than less well is kinder.

And it is immediately actionable, which is a property this series has repeatedly found in claims that outlive their evidence. A questionnaire and two versions of a slide deck is a plan.

Against which, the honest alternative is harder to sell. Practices with an evidence base tend to be effortful rather than clever, which is the next section.

Not An Argument Against Tailoring

Because this could be misread as licence to teach everyone identically. Ours.

Four observations.

Prior knowledge plainly should change how you train somebody, and nothing here disputes that. A new hire and a ten-year veteran need different things.

So does the material. Some things are inherently visual and some are inherently verbal, and matching the method to the content is a different claim from matching it to the person.

The review itself distinguishes related literatures, its headings naming aptitude-by-treatment and personality-by-treatment interactions as separate bodies of work[3], which we did not obtain and which are not the same question.

And the claim under examination is narrow and specific. That a self-reported presentation preference should determine the format of instruction, which is one idea among several and the one with the missing evidence.

What The Review Recommends Instead

The constructive half of the conclusion, which is rarely quoted alongside the negative half.

It states that "limited education resources would better be devoted to adopting other educational practices that have a strong evidence base, of which there are an increasing number."[2]

Four observations, ours.

"Of which there are an increasing number" is an optimistic clause in an otherwise negative conclusion, and it is the authors' own.

We did not obtain the review's account of what those practices are, and will not list them from memory, which would be exactly the failure this series documents.

Two of the four authors are known for work on the spacing and testing of learning, on our reading of the biographical notes we obtained[3], which suggests where their answer would point without us asserting it.

And the framing is the useful part regardless of the content. The question is not whether learning styles are real but what else the same budget could buy, which is a question a business is well equipped to ask.

Your Training Budget

The application. Ours, and not training advice.

Four points.

If a training provider assesses learning styles, ask what evidence the matching rests on. The answer to look for is a study with all four cells, and the review's finding is that very few exist.

Ask what the same money would buy in a practice with an evidence base, which is the review's own recommendation and reframes the decision usefully.

If you want to test it yourself, the design is expensive and the arithmetic is not. Two groups, two versions, four cells, and multiply the row differences.

And be honest about the sample problem. A firm with forty staff cannot run this study, on the power figures the eighty-eighth article computed, which means the practical answer is to rely on published evidence rather than to generate your own.

Preferring And Benefiting Are Different

The gap the instruments cannot bridge. Ours.

Four observations.

The review's abstract records that assessments typically ask people to evaluate what sort of information presentation they prefer[1], which is a report about liking.

Liking a format and learning from it are separable and frequently opposed. A method that feels easy can produce less durable learning than one that feels effortful, which is a well-known theme in the memory literature two of these authors work in.

If that holds, a preference instrument may be systematically pointed at the wrong thing, selecting for comfort rather than for benefit.

And we flag that as our own inference rather than the review's finding. We did not obtain the review's treatment of this, and are reasoning from a single word in its abstract and from the authors' known field.

Which is a caution we would apply to ourselves as much as to anyone. Reasoning from an author's known interests to what they must have argued is exactly the move this series criticises, and we have flagged it rather than presented it as established.

What A Good Answer Sounds Like

Because "ask for the evidence" is easy to say and harder to act on. Ours.

Four observations.

A provider saying "our participants who took the visual track scored higher afterwards" has described the weakest design in the table above, which finds improvement in every scenario.

One saying "visual learners did better with our visual materials than with our verbal ones" has described the second design, which gives the same answer whether matching helps or not.

A good answer names the fourth cell without being prompted. "We also taught verbal learners with the visual materials, and here is what happened" is the sentence that indicates somebody understands the question.

And an honest answer is also available and perfectly respectable. "We have not tested it and here is why we think the course is worth your money anyway" is a defensible position, and considerably better than a study that cannot bear the weight.

One reason to prefer that answer, ours. A provider who overclaims about the evidence will overclaim about the outcome, and the first is checkable before you buy while the second is not.

Where Else This Test Applies

The reason to keep the test rather than the topic. Ours.

Four observations.

The crossover requirement applies to any claim that a treatment should be matched to a type, which is a very common shape in business.

Examples, ours and untested. This sales script works better for this customer segment. This management style suits this personality. This payment structure motivates this kind of employee. Each is a meshing claim.

And each is usually supported by the partial design. We tried it with that segment and it worked, which cannot distinguish a genuinely segment-specific effect from something that works better on everybody.

So the transferable question is a single one. Did you also try it on the other group? If not, you have learned that the thing works, which is worth knowing, and nothing about whether to match.

One reason to keep asking it even when the answer is inconvenient, ours. Segmenting has real costs: two sets of materials, a classification step, and the risk of putting people in the wrong bucket.

Which means the default should be the single best method for everyone, and the burden sits with the case for splitting, rather than with the case for treating people the same.

Bibliographic Note

The series keeps a count, and this article produced two.

The review's page range is given as 105–119 by the database record[1] and as 106–119 in one circulating citation.

And the article is dated December 2008 while its assigned identifier encodes 2009[1], so a reader searching by year may look in the wrong volume.

Three observations, ours.

The year discrepancy is nobody's error, arising from an article published in December and assigned an identifier in the following year, which is routine.

It is nonetheless the kind of thing that sends a search astray, which is the practical test we have applied throughout.

That brings the running count of bibliographic variants across this series to thirty-six.

One pattern across all thirty-six is now worth stating, ours. Not one of them has changed a finding. Every variant we have recorded would mislead a search and none would mislead a reader who found the paper.

Which is a mildly reassuring result from a count we began expecting to be more damning. Citation practice drifts and substance does not, at least across the sample this series happens to have looked at.

What To Do

Ask for all four cells. A claim that a method suits a type requires evidence about both types under both methods, and anything less cannot answer it.

Multiply the two row differences and check the sign. Negative means a crossover; positive means one method is simply better, however large the interaction.

Do not accept "it worked for that group" as evidence for matching. On our own arithmetic that result is identical whether the hypothesis is true or false.

Watch for the ordinal case specifically. A real interaction where one method still wins for everyone is the pattern most likely to be reported as support.

Consider that matching can make things worse. On our own invented figures it cost five points by routing half the group to the method that was worse for everybody.

Ask what else the budget would buy. The review's recommendation is to redirect resources to practices with a strong evidence base, not to stop training.

Do not overclaim in the other direction. The authors state it would be an error to conclude that all versions have been tested and found wanting, and many simply have not been tested.

And apply the test beyond this topic. Any claim that a treatment should be matched to a type has this shape, and the question is always whether you also tried it on the other group.

The Limits Of This Analysis

Several caveats matter. This article discusses research on instruction and is not training or educational advice; the applications are our own reasoning and untested. Everything is verified to August 2026. We did not obtain the full review, only its abstract from a database, its conclusion from a repository record, and its table of contents with some body passages from a copy hosted on an academic networking site; we therefore cannot say how many studies it examined, which ones it judged to have used appropriate methods, or what those studies found in detail, and our account of "several" contradictory results is the review's own word rather than a count we verified. We did not obtain the review's account of which alternative practices have a strong evidence base, and have deliberately not listed any from memory. The three conditions for a valid test and the description of the review as commissioned reach us through two non-academic sources, flagged at every use. We obtained none of the subsequent literature, which appears here only as citations in a reference list. The paper's distinction between the full meshing hypothesis and a weaker version requiring only the crossover interaction is a subtlety we obtained in fragment and did not fully resolve, and a reader relying on that distinction should obtain the paper. All arithmetic is ours and every score in it is invented: the three scenarios are constructions demonstrating what the designs can and cannot distinguish, not estimates of anything real, and the five-point cost of matching depends entirely on the first scenario being the true one, which nobody knows. And this article contains one withdrawn section: we first computed a false-confirmation rate that was circular by construction, and replaced it with a comparison of what four designs can distinguish.

Frequently Asked Questions

What did the 2008 review conclude?
That although the literature is enormous, very few studies used an experimental methodology capable of testing the claim, and of those that did, several found results that flatly contradict the meshing hypothesis. It concludes there is no adequate evidence base to justify incorporating learning-styles assessments into general educational practice.
Does that mean learning styles do not exist?
The authors explicitly decline to say so. Their conclusion states it would be an error to conclude that all possible versions have been tested and found wanting, since many have simply not been tested. The finding is that the evidence is missing, not that it is negative.
What is a crossover interaction?
A pattern where each group does better with its own matched method, so the two lines cross. Visual learners score higher with visual instruction and verbal learners score higher with verbal instruction. Only that pattern supports matching, because it means neither method is better in general.
How do I test for it?
Take the difference between the two methods for the first group, take it for the second group, and multiply. A negative product means a crossover. On our own invented scenarios that gives plus 100 where the hypothesis is false, minus 100 where it is true, and plus 150 for an interaction that still does not support matching.
Why is "it worked for that group" not enough?
Because on our own arithmetic that result appears identically whether the hypothesis is true or false. If visual teaching is simply better for everyone, visual learners will do better with it, and so will verbal learners. The study never looks at the second row, so it cannot tell the cases apart.
Could matching actually hurt?
On our own invented figures, yes. If one method is better for everyone, matching routes half your people to the worse method and costs five points on a hundred-point scale, plus the assessment cost and the cost of building two versions of everything. That depends on a scenario nobody has established.
What should I do with a training budget instead?
The review's own recommendation is that limited resources would be better devoted to practices with a strong evidence base, of which it says there is an increasing number. We did not obtain its account of which those are, and have deliberately not listed any from memory.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article withdrew one of its own sections as circular and replaced it with a test readers can apply themselves.

References

  1. National medical library database record for Pashler, H., McDaniel, M., Rohrer, D., & Bjork, R. (2008), Learning Styles: Concepts and Evidence, Psychological Science in the Public Interest, 9(3), 105–119, December, DOI 10.1111/j.1539-6053.2009.01038.x, reproducing the abstract: that the term learning styles refers to the concept that individuals differ in regard to what mode of instruction or study is most effective for them; that proponents of learning-style assessment contend optimal instruction requires diagnosing individuals' learning style and tailoring instruction accordingly; that assessments typically ask people to evaluate what sort of information presentation they prefer, such as words versus pictures versus speech, and what kind of mental activity they find most engaging, although assessment instruments are extremely diverse; and that the most common but not the only hypothesis about instructional relevance is the meshing hypothesis, according to which instruction is best provided in a format that matches the preferences of the learner. Note: a national medical library database record. Our source for the abstract. Recorded also as a bibliographic variant: the record dates the article December 2008 while its assigned identifier encodes 2009. pubmed.ncbi.nlm.nih.gov
  2. University repository record for the same review, reproducing its conclusion: that although the literature on learning styles is enormous, very few studies have even used an experimental methodology capable of testing the validity of learning styles applied to education; that of those which did use an appropriate method, several found results that flatly contradict the popular meshing hypothesis; that the authors conclude there is at present no adequate evidence base to justify incorporating learning-styles assessments into general educational practice; that limited education resources would better be devoted to adopting other educational practices that have a strong evidence base, of which there are an increasing number; and that however, given the lack of methodologically sound studies, it would be an error to conclude that all possible versions of learning styles have been tested and found wanting, since many have simply not been tested. Note: a university repository record. Our source for the conclusion and for the caveat, which we have quoted in full because it is routinely omitted from accounts of this review. digitalcommons.usf.edu
  3. Copy of the review hosted on an academic networking site, reproducing its table of contents, which includes the section headings "Interactions as the Key Test of the Learning-Styles Hypothesis," "Style-by-Treatment Interactions: The Core Evidence Is Missing," "Learning-Styles Studies With Appropriate Methods and Negative Results," "Aptitude-by-Treatment Interactions," "Personality-by-Treatment Interactions," and "Costs and Benefits of Educational Interventions"; and reproducing body passages referring to the style-by-method crossover interaction, including a statement that a weaker version of the hypothesis requires only that crossover interaction and does not require that the optimal method for each group match that group's learning style, and a passage on dividing subjects into groups of learners with high auditory ability and high visual ability. It also carries the authors' biographical notes. Note: an academic networking site hosting the paper itself. Our source for the section headings and for the crossover terminology. We did not obtain the complete text, and the distinction between the full meshing hypothesis and the weaker crossover-only version reaches us in fragment. researchgate.net
  4. Education advocacy organisation's summary of the review, recording that the authors pointed out that experiments designed to investigate the meshing hypothesis would have to satisfy three conditions, the first being that learners would be allocated to two or more groups, such as visual and auditory, on the basis of some assessment of their presumed learning style; and citing the review as Psychological Science in the Public Interest, 9(3), 105–119. Note: an advocacy organisation, NOT an academic source, flagged at every use. Our source for the three stated conditions, of which our reproduction gives only the first in full. deansforimpact.org
  5. Professional council's article on the topic, recording that the review was commissioned by a professional psychological association; that any credible validation of learning-style-based instruction must demonstrate a particular type of statistical result known as a crossover interaction; and that the required experimental design involves several necessary criteria, beginning with participants being assessed and classified based on their purported learning style. Its reference list names Kirschner, P. A. (2017), Stop propagating the learning styles myth, Computers & Education, 106, 166–171; Knoll, A., Otani, H., Skeel, R., and Horn, K. (2017), Learning style, judgements of learning, and learning of verbal and visual information, British Journal of Psychology, 108, 544–563; Rohrer, D., and Pashler, H. (2012), Learning Styles: Where's the Evidence?, 46(7), 634–635; and Marwaha, K., and Sharma, U. (2025), Debunking Learning Styles: Analyzing Key Predictors of Academic Success in Dental Education, Advances in Physiology Education. Note: a professional council's article, NOT a peer-reviewed source, flagged at every use. Our source for the crossover requirement in plain terms and for the description of the review as commissioned, neither of which we verified independently. The subsequent literature named here reaches us as citations only and we obtained no part of any of it. gc-bs.org

This article discusses research on instruction and is not training or educational advice. The full review was not obtained, only its abstract, conclusion, table of contents and some body passages; the number of studies it examined and their identities are therefore not reported here. Two sources are non-academic and are flagged at every use. All arithmetic is the authors' own and every score in it is invented. This article contains one withdrawn section.