Artificial WastelandAt Full Strength · psychology

Nine Points for Ten Minutes

In 1993 a one-page letter in Nature reported that 36 college students scored 3.56 standard-score points higher on Stanford-Binet spatial subtests after ten minutes of a Mozart sonata than after silence, and translated the gap as 8 to 9 IQ-equivalent points. This page rebuilds that claim from its printed numbers, then plants the 1993 gap into the printed tables of the 1999 replication that tested it: Western Ontario’s test would have jumped from F = 2.05 (printed 1.99) to 26.2, while Montreal’s arm alone was too weak to settle it. As measured, the two arms put the gap at 0.49 of those points (standard error 0.79).

The 1993 gap, and the 1999 replication that tested it

Mozart-minus-silence gaps in Stanford-Binet standard age score (SAS) points, with 95% intervals: the 1993 claim on top (its interval is the page’s arithmetic on the printed t), then the 1999 replication’s two SAS arms and the page’s pooling of them.

As measured: Montreal −1.75, Western Ontario +1.31, the two pooled +0.49 SAS points. Western Ontario’s own test reads F = 2.0; Montreal’s t = −1.14.

Mozart-minus-silence gaps: the 1993 claim, the two 1999 arms, and their poolingA dot and interval for each: the 1993 gap near 3.6 points, Montreal below zero, Western Ontario a little above zero, and the pooled control near half a point.

Everything below is computed in your browser from four small files this page froze from the printed record. A number taken from a publication is marked as printed and cited where it appears; every P value, interval, power and pooled estimate that is not so marked is the page’s own. The check at the bottom recomputes the page against itself while you read it.

I · the claim, at full strength

Three means, one table, and eight or nine points

On 14 October 1993 Nature printed a one-page letter from Frances H. Rauscher, Gordon L. Shaw and Katherine N. Ky of the Center for the Neurobiology of Learning and Memory at the University of California, Irvine. 36 college students each sat three ten-minute listening conditions: Mozart’s sonata for two pianos in D major (K. 448, as Rauscher’s 1999 reply names it), a tape of relaxation instructions, and silence. Each condition was followed by one of three Stanford-Binet subtests: pattern analysis, matrices, and paper folding and cutting. The letter reports the result in a bar chart and one paragraph:

The music condition yielded a mean SAS of 57.56; the mean SAS for the relaxation condition was 54.61 and the mean score for the silent condition was 54.00. To assess the impact of these scores, we ‘translated’ them to spatial IQ scores of 119, 111 and 110, respectively. Thus, the IQs of subjects participating in the music condition were 89 points above their IQ scores in the other two conditions.

F. H. Rauscher, G. L. Shaw and K. N. Ky, “Music and spatial task performance”, Nature 365, 611 (1993), doi:10.1038/365611a0. The byline prints Katherine N. Ky; Crossref and the Nature record give Catherine N. Ky.

That is the claim at the strength its authors printed it, and it is narrower than its fame. It is about a gap on spatial reasoning subtests in college students, and the letter itself says how long it lasted:

The enhancing effect of the music condition is temporal, and does not extend beyond the 10–15-minute period during which subjects were engaged in each spatial task.

Rauscher, Shaw and Ky 1993, p. 611

The 1993 figure, redrawn from its printed numbers

Mean SAS after music, relaxation and silence, with the IQ-equivalent axisThree bars at 57.56, 54.61 and 54.00 on a SAS axis from 50 to 60, with a right-hand IQ-equivalent axis that follows the translation chosen below.

The printed figure has two axes on the same gridlines: SAS 50 to 60 on the left and IQ equivalents 100 to 125 on the right, so the claimants’ own drawing fixes a line of 2.5 IQ-equivalent points per SAS point. The labels above each bar give the IQ equivalent under the translation chosen in the panel further down, and the integer the letter printed.

The significance, reproduced

A one-factor (listening condition) repeated measures analysis of variance (ANOVA) performed on SAS revealed that subjects performed better on the abstract/spatial reasoning tests after listening to Mozart than after listening to either the relaxation tape or to nothing (F2,35 = 7.08; P = 0.002). The music condition differed significantly from both the relaxation and the silence conditions (Scheffe’s t = 3.41, P = 0.002; t = 3.67, P = 0.0008, two-tailed, respectively). The relaxation and silence conditions did not differ (t = 0.795; P = 0.432, two-tailed).

Rauscher, Shaw and Ky 1993, p. 611 (the scan prints Scheffe without an accent)

The letter prints no degrees of freedom for its t values. A paired comparison of 36 people has 35, and with that one inference each printed t gives back its printed P:

Each printed t through the page’s two-tailed t distribution with 35 degrees of freedom
comparisonprinted tprinted Pthe page’s PSD of the gap it implies
music against relaxation3.410.0020.001655.19
music against silence3.670.00080.0008015.82
relaxation against silence0.7950.4320.4324.60

Each recomputed P rounds to the printed one, so the claim’s significance reproduces at full strength. The printed gap of 3.56 SAS points over silence with t = 3.67 means the students’ individual music-minus-silence gaps had a standard deviation of about 5.82 points, a within-person effect of dz = 0.61, and a 95% interval for the gap of 1.59 to 5.53 SAS points (the page’s arithmetic on the printed numbers).

The letter calls the first of these “Scheffe’s t”, but the P values it prints are the plain two-tailed values for each t, as the table shows. A Scheffé criterion for three conditions refers t² / 2 to an F distribution with 2 and 35 degrees of freedom (or 70, pooling the error across conditions), and would give 0.0066 or 0.0046 for t = 3.41, and 0.0034 or 0.0021 for t = 3.67. Both music contrasts stay below 0.01 on either reading, so the claim’s significance does not turn on which one is meant.

The omnibus F is the one printed statistic the page cannot recompute: it needs each student’s three scores, which were never published. From the three printed paired t values the page can rebuild the covariance of the within-person differences, and from that two versions of the omnibus test: a repeated-measures F that assumes sphericity, 9.55 (9.44 to 9.65 over the rounding of the six printed numbers), and a multivariate Hotelling F, 7.40 (7.32 to 7.48). Both are at least as strong as the printed 7.08, so the page carries the claim at the printed strength.

The claimants also tested the most obvious rival explanation themselves:

Pulse rates were taken before and after each listening condition. A two-factor (listening condition and time of pulse measure) repeated measures ANOVA revealed no interaction or main effects for pulse, thereby excluding arousal as an obvious cause.

Rauscher, Shaw and Ky 1993, p. 611

What an IQ-equivalent point is

The famous number was never measured on an IQ scale. It is a translation of subtest scores, and the letter documents it:

These were then converted to SAS using the Stanford–Binet’s SAS conversion table of normalized standard scores with a mean set at 50 and a standard deviation of 8. IQ equivalents were calculated by first multiplying each SAS by 3 (the number of subtests required by the Stanford–Binet for calculating IQs). We then used their area score conversion table, designed to have a mean of 100 and a standard deviation of 16, to obtain SAS IQ equivalents.

Rauscher, Shaw and Ky 1993, p. 611, “Scoring”

On the figure’s own axis the three means become 118.9, 111.5 and 110.0, each within 0.6 of the printed integers, and the gaps become 7.4 over relaxation and 8.9 over silence. The printed 8 and 9 are differences of the rounded equivalents (119 minus 111, and 119 minus 110). That is a rounding gap, and it is the one place where the page’s reproduction of the claim is weaker than printed: 7.4 before rounding, 8 after. The page prints both.

Why 2.5 points per SAS point, when a plain change of scale from one subtest (SD 8) to an IQ scale (SD 16) gives 2.0? Because of the multiply-by-3 step. The area table converts a sum of subtest scores, and a sum of three correlated subtests spreads out less than three times one subtest. If the three have an average correlation r, their sum has an SD of 8 √(3 + 6r), and the slope becomes 6 / √(3 + 6r). For a line through SAS 50 = IQ 100, the printed integers need a slope between 2.447 and 2.495, which in this model means r between 0.464 and 0.502. That r is fitted by the page to the printed values, not published anywhere; the table itself is not public. The letter does say the three tasks “correlated at the 0.01 level of significance” in its sample.

Translate the gap yourself: the page’s linear model of a table that is not public

Translation
IQ-equivalent points per SAS point2.50
music over relaxation, over silence7.4 and 8.9
music, relaxation, silence118.9, 111.5, 110.0

With r = 0 and three subtests the same 3.56 SAS points become 12.3 points over silence and 10.2 over relaxation; with one subtest, 7.1 and 5.9. The model refuses a number of subtests other than 1, 2 or 3, and an r outside 0 to 1.

The 1999 control translated too, and printed its own IQ equivalents. Its Montreal Mozart group averaged 57.31 SAS, which the 1993 axis would call 118.3; Steele and colleagues printed 114, 4.3 points lower for nearly the same subtest score. Their two printed gaps each work out to 2.29 points per SAS point (the page’s arithmetic on rounded integers), and their letter does not say which table or how many subtests they used. That spread is the size of the uncertainty an “IQ-equivalent” carries. The page therefore plants and tests the claim in SAS points, the unit both groups measured in.

The claimants’ own first-task data

Chabris’s 1999 table (Act II) includes rows computed from data the claimants supplied, using only the first task each student sat, so that no listening condition could carry over. In those rows the largest effect was on paper folding and cutting: d = 1.389 against silence and 1.622 against relaxation, with 8 students per comparison and printed P = 0.140 and 0.094, so small samples and neither significant on its own, against 0.097 and 0.000 for matrices and 0.289 and −0.685 for pattern analysis. Paper folding and cutting is exactly the task the 1999 control used.

II · the control

Prelude or requiem: Nature, 26 August 1999

Six years later Nature printed three items under one title, “Prelude or requiem for the ‘Mozart effect’?”: a replication run in three laboratories by Kenneth M. Steele and eight co-authors, a meta-analysis by Christopher F. Chabris, and a reply by Frances Rauscher. The replication is the control this page tests. It did not settle the question on its own: it was answered in the same issue, and a supportive synthesis followed in 2000, so the verdict below dates the decision to the 2010 meta-analysis. The replication’s authors wrote that “The experimental designs replicated the original study at the University of Montreal (UM); other standard designs were used at the Appalachian State University (ASU) and the University of Western Ontario (UWO)”, and that “SAS values in the UM and UWO studies are quite similar to the original report, indicating that the subjects had similar intellectual skills.” Their result:

The results show that listening to the Mozart sonata produced no differential improvement in spatial reasoning in any experiment.

K. M. Steele, S. Dalla Bella, I. Peretz, T. Dunlop, L. A. Dawe, G. K. Humphrey, R. A. Shannon, J. L. Kirby Jr and C. G. Olmstead, Nature 400, 827 (1999), doi:10.1038/23611
Steele et al. 1999, Table 1: Effect of listening condition on scores from the Paper Folding and Cutting task (transcribed from the printed table)
listening conditionmeans.e.N as printed
UM, Stanford–Binet SAS scores
Mozart57.311.2632
Silence59.060.8832
UWO, Stanford–Binet SAS scores
Mozart55.580.6424
Silence54.270.6721
Relaxation music54.140.3122
ASU, number correct from 16 items
Mozart10.780.7418
Silence11.420.9117
Minimalist music10.830.7818
10 min relaxation10.890.8618
20 min relaxation11.070.9815

The page recomputes each laboratory’s own test from these rows. Montreal: t(30) = −1.139, P = 0.264, against the printed t(30) = 1.14, P = 0.263. Western Ontario: F(2, 64) = 2.047, P = 0.137, against the printed 1.99, P = 0.145. Montreal’s t matches the printed digits, and its P (0.264 against the printed 0.263) is inside the 0.259 to 0.269 that the rounding of the table allows. Western Ontario’s F does not match the printed digits, and need not: over the rounding of the printed means and standard errors the Montreal t can be anywhere from 1.127 to 1.150 and the Western Ontario F from 1.984 to 2.113, and both printed values sit inside. That validates the recomputation the page is about to reuse.

Appalachian State’s printed F(4, 81) = 0.33, P = 0.86 is a treatment effect in a pretest and posttest design and cannot be recomputed from the posttest table. The page prints it as printed; a posttest-only one-way F, the page’s own figure, is 0.094 (P = 0.98).

Montreal’s N column. The table prints 32 on both Montreal rows. Each participant sat paper folding and cutting once, after either Mozart or silence (the other treatment preceded the matrices task), so the page reads 32 participants in all, 16 per condition. The printed df of 30 fits that reading, and so does Chabris’s first-task row for this laboratory (N = 16). The printed mean d below fits it too, but only with Chabris’s value for Appalachian State, so it is not independent evidence; the df and Chabris’s row carry the reading on their own. Read as 32 per condition, the df would be 62 and the mean d 0.042. The t itself needs only the means and standard errors.

The letter summarised all three experiments in one number:

Conversion of the Mozart and silence comparisons into a measure of effect size indicated that the music had little impact (mean d = 0.003). A requiem may therefore be in order.

Steele et al. 1999, p. 827

The page reproduces that mean as 0.0028, from Montreal’s d = −0.403, Western Ontario’s 0.422 and, for Appalachian State, the −0.011 that Chabris’s table prints for this laboratory. Which ASU contrast Steele and colleagues converted is not printed. With the page’s own posttest value for ASU, −0.186, the mean would be −0.055.

Appalachian State’s value in the mean d

Mean d of the three laboratories: 0.0028 (printed 0.003).

The honest reading of the table is not “zero everywhere”. Western Ontario’s Mozart group scored +1.31 SAS points above silence (t = 1.41, P = 0.17), in the claimed direction and 37% of its size; Montreal’s scored −1.75. Pooled by inverse variance (the page’s own step), the two arms scored in the claim’s unit put the Mozart-minus-silence gap at 0.49 SAS points (standard error 0.79), 95% interval −1.06 to 2.05. On the claimants’ own axis that is 1.2 IQ-equivalent points (−2.7 to 5.1) against the claimed 8.9.

Chabris’s meta-analysis, the same week

Meta-analysis combining the effect sizes reported for all 20 published Mozart-to-silence comparisons (Table 1, top), involving a total of 714 subjects, yields an average cognitive enhancement of d = 0.09 standard deviations, or only 1.4 IQ points.

C. F. Chabris, Nature 400, 826–827 (1999), doi:10.1038/23608

Chabris printed every row he pooled, so the page can pool them again. He weighted each d by its degrees of freedom (N − 2) and combined P values as Z scores. On his 28 rows the page gets 0.092 for all silence comparisons (printed 0.09), −0.040 for abstract reasoning (printed −0.04), 0.140 for spatial–temporal tasks (printed 0.14) with Z = 1.131 and P = 0.258 (printed Z = 1.14, P = 0.26; the rounding of his printed P values moves the page’s Z between 1.124 and 1.139), and against relaxation instructions 0.204 overall and 0.562 for spatial–temporal tasks (printed 0.20 and 0.56). At an IQ standard deviation of 15, the value that reproduces his conversions, those are 1.4 and 2.1 IQ points. His explanation was a small “‘enjoyment arousal’ effect of Mozart’s music on difficult spatial tasks”.

Pool Chabris’s rows

Compared with
Task class
His row for Western Ontario (N = 45)

d = 0.092 over 20 comparisons and 714 people; Z = 0.77, P = 0.44; 1.4 IQ points at SD 15

Chabris 1999, Table 1, all 28 rows as printed (transcribed; the reference numbers are listed below the table). Bold: the claimants’ own first-task rows. Italic, with “no” in the first column: rows outside the pool you chose.
in this poolagainsttaskref.NdPclass
The table is drawn from the frozen file when the page’s scripts run.
Chabris’s reference numbers

    Chabris named the two classes by example. Putting maze completion, pattern analysis and short-term memory in neither is the page’s inference; it reproduces both printed class values. In the PDF’s text layer every minus sign reads as a leading 1 (−0.989 reads 10.989), so every value here was read against the rendered table.

    A coding check that moves the critics’ number toward the claim. Chabris’s row for Western Ontario prints d = 0.017, P = 0.956 for 45 people. Steele’s printed Western Ontario means give d = 0.42, P = 0.17. The page could not reproduce Chabris’s value from the printed source. His table marks with an asterisk the rows for which he obtained extra information from the authors; this row carries none (his Montreal row does), so the page cannot say what he computed it from. With the recomputed row his silence d becomes 0.118 and his spatial–temporal d 0.175 (Z = 1.53): larger, and still not statistically significant.

    The claimants’ reply, in the same issue

    The comments by Chabris and Steele et al. echo the most common of these: that listening to Mozart enhances intelligence. We made no such claim. The effect is limited to spatial–temporal tasks involving mental imagery and temporal ordering.

    Steele et al. find no Mozart effect in three differently designed studies. Not one design replicated the original reports, and they introduced several methodological concerns. For example, spatial–temporal task performance varies widely between individuals, making randomization an inefficient way to ensure uniform before-treatment task proficiency. What measures were taken by the two studies using between-subjects designs to tackle this? Was testing done blind, as in other replications …?

    Although the Mozart effect cannot be found under all laboratory conditions, as discussed by Steele et al., several studies have successfully replicated it … Because some people cannot get bread to rise does not negate the existence of a ‘yeast effect’.

    Frances H. Rauscher, “Rauscher replies”, Nature 400, 827–828 (1999)

    The reply’s evidence, as it gives it: studies it counts as replications (its refs 1 to 8, 11, 13 and 14, and manuscripts in preparation); the statement that “The effect works for not just one spatial–temporal task, as claimed by Chabris, but for three” (its refs 5 and 8 and a manuscript in preparation); rats exposed to the sonata in utero and for 60 days after birth that learned a spatial maze faster; students who showed the effect after Mozart though they rated Mendelssohn as more arousing; and a report that the sonata reversed epileptiform activity in comatose patients while control music did not. The page tests the randomization objection in Act III. On blinding, the printed 1999 record does not say whether the testing was blind.

    III · the control on the control

    Could the 1999 control have seen 3.56 points?

    A control that could not have found the effect proves nothing by failing to find it. So the page gives the claim its printed size and plants it. Grade A (injection), at the level of the printed summary table. The recomputed Montreal t and Western Ontario F depend on the data only through each group’s mean, standard deviation and N. Adding the claimed gap to every participant in a Mozart group moves that group’s mean by exactly the gap and leaves its standard deviation unchanged, so planting the gap into the printed Mozart row and rerunning the same recomputation is exactly what planting it into the unpublished rows would do. The assumption it rests on is printed here: an additive, constant effect with no ceiling. The recomputation was validated first, above, against the printed t and F.

    The size planted is 3.56 SAS points, the claim’s own music-minus-silence gap, computed from the claim file, in the claim’s own unit, on the claim’s own test family. No conversion to d and no scale mapping is needed, because Montreal and Western Ontario scored the same Stanford-Binet subtest in SAS. It does not overstate the claim for this task: it is an average over three subtests; the claimants’ reply says the effect works for more than one spatial–temporal task, and their own first-task rows, small as they are, put the largest d on paper folding and cutting. So the claim gives no reason to expect less than 3.56 on this task, and anything the control sees at 3.56 it sees at more.

    Plant the claim into the control’s own tables

    Claimed gap to plant
    Significance level
    Arms in the pooled estimate

    Each arm’s own test, as measured and with 3.56 SAS points planted into its Mozart row; power at that size and level 0.05. The pooled row gives the gap ± one standard error.

    armfound?poweras measuredclaim planted
    Montreal, tnot recovered0.61t = −1.14, P = 0.26t = +1.18, P = 0.25
    Western Ontario, F (found?) and t (power)recovered0.96F = 2.05, P = 0.14; gap t = +1.41F = 26.2, P = 4.9 × 10−9; gap t = +5.25
    the two pooledrecovered0.99+0.49 ± 0.79, z = +0.62+4.05 ± 0.79, z = +5.10

    The control could have confirmed the claim: planted at 3.56 points, it is recovered by the two Stanford-Binet arms pooled, with power 0.99.

    Power of each arm against the planted gapPower rises with the planted gap; Western Ontario and the pooled estimate pass 0.8 well below the claimed 3.56, Montreal does not reach it by 3.56.
    MontrealWestern Ontario (Mozart against silence)pooledthe claimed sizes

    Power here is the page’s: each arm’s Mozart-against-silence test (a two-sided t with that arm’s own standard error and degrees of freedom, by numerical integration), and a normal test for the pooled estimate. For Western Ontario, “found?” uses the laboratory’s own three-group F, as its authors tested it, and “power” is for its Mozart-against-silence t; the three-group F’s own power at 3.56 and 0.05 is 0.997. The page’s rule for the sentence above: the planted gap is recovered at the chosen significance level by the chosen arms, and their power at that size is at least 0.80. That line is the conventional minimum for a planned test; at the default setting the verdict does not lean on it, since the pooled power is 0.99. Planting zero gives back the as-measured numbers exactly.

    At the claimed 3.56 points, Western Ontario’s F rises from 2.05 to 26.2 (P = 4.9 × 10−9) and its Mozart-versus-silence t from 1.41 to 5.25: recovered, with power 0.96 for the gap itself (0.997 for its own three-group F). Pooled with Montreal the gap goes to 4.05 (z = 5.1), power 0.99. The control could have confirmed the claim: planted at 3.56 points, it is recovered by the two Stanford-Binet arms pooled, with power 0.99.

    Montreal is different, and the page says so plainly: its t moves only from −1.139 to +1.18 (P = 0.25). The planted claim is not found there, and its power at the claimed size was 0.61. Montreal’s arm on its own could not have confirmed the claim under the page’s rule: in the planted table the claim does not reach significance, and a test with that power would miss a true gap of that size with probability 0.39. So its null is not evidence against the claim by itself. The verdict of this act rests on Western Ontario and on the page’s pooling of the two arms. At the smaller relaxation gap the powers are 0.46, 0.87 and 0.96. The claimed 3.56 lies 3.9 standard errors above the pooled measured gap.

    This is the page’s answer to the claimants’ objection that between-subject randomization is “an inefficient way to ensure uniform before-treatment task proficiency”. The inefficiency is real and visible: it is in the standard errors, and it is why Montreal’s arm alone was underpowered. It was not enough to hide the claimed gap from Western Ontario, whose students varied less. That is the page’s reading of the numbers, not a concession the claimants made.

    Appalachian State stays out of the power statement. It reported the number correct out of 16 items, with no printed conversion to SAS, so planting 3.56 SAS points there would need a scale mapping nobody printed. The button in the first panel asks for it and is refused. The Psychological Science replication by Steele, Bass and Crook (July 1999, 125 participants, raw item counts, treatment F(2, 122) = 0.11 as printed) stays out for the same reason.

    The claimants’ method on data with nothing in it

    Not applicable here, and the page says why. The record did not turn on the claimants’ analysis: the page reproduces their printed P values from their printed t statistics, and the claim was set aside on replication, not on a fault found in how the 36 students’ paired scores were tested. For the record, on data with no effect that paired test gives a music-minus-silence t of 3.67 or more with probability 0.0004, about 4 times in 10,000: half the printed two-tailed P. The page states this arithmetic rather than simulating it. The “method on nothing” the history does turn on belongs to the critics’ side, and it is in Act V.

    IV · the verdict

    Dissolved, and what survives

    DISSOLVED

    as of · decided by independent replication · 17 years from the claim to the 2010 meta-analysis

    Scoped to the claim as printed in 1993: that ten minutes of the Mozart sonata K. 448 raised Stanford-Binet spatial reasoning scores by 3.56 standard-age-score points over silence (2.95 over relaxation instructions), translated as 8 to 9 IQ-equivalent points. Filed as DISSOLVED, not ARTEFACT. Fudin and Lembessis (2004) raised questions about the 1993 scoring, experimental design, IQ measure and statistical analyses (the page has read their abstract, not the paper), but no instrument, analysis, contamination or perception fault in the 1993 experiment was established in the record, and the page reproduces the claimants’ own significance from their own printed numbers. The 2023 sentence quoted below attributes the phenomenon to low study power and bias-related measurement artifacts in the literature, not to a named fault in the 1993 experiment. What removed the claim was more data: an independent three-laboratory replication and two meta-analyses (1999 and 2010), with a preliminary 2025 synthesis pointing the same way and one supportive 2000 synthesis on the other side.

    A small difference between music and silence before a spatial task, not specific to Mozart, survives in the 2010 meta-analysis (d = 0.37 for Mozart against no music and d = 0.38 for other music against no music, both before its downward correction for publication bias). The page does not call that zero.

    Sources that establish it

    1. Jakob Pietschnig, Martin Voracek, Anton K. Formann, 2010, Mozart effect–Shmozart effect: A meta-analysis, Intelligence 38 (3), 314-323, DOI 10.1016/j.intell.2010.03.001 (abstract quoted from ERIC record EJ882611) doi.org/10.1016/j.intell.2010.03.001
    2. Sandra Oberleiter, Jakob Pietschnig, 2023, Unfounded authority, underpowered studies, and non-transparent reporting perpetuate the Mozart effect myth: a multiverse meta-analysis, Scientific Reports 13, 3175, DOI 10.1038/s41598-023-30206-w (CC BY 4.0; cited only for its dated sentence on the 1993 spatial claim, since the paper itself is about epilepsy) doi.org/10.1038/s41598-023-30206-w
    3. Christopher F. Chabris, 1999, Prelude or requiem for the ‘Mozart effect’?, Nature 400, 826-827, DOI 10.1038/23608; Kenneth M. Steele and eight co-authors, 1999, same title, Nature 400, 827, DOI 10.1038/23611 (the 1999 milestone) doi.org/10.1038/23611
    4. Marie Meunier, Ezio Tirelli, Nancy Durieux, François Léonard, 2025, Mozart effect: A meta-research study on statistical power, effect size, and false discovery rate: Preliminary results, poster, ORBi 2268/334434 (a preliminary poster, not a journal article) orbi.uliege.be/handle/2268/334434

    What would change it

    A preregistered, blind, multi-laboratory replication that includes the originating laboratory, runs K. 448 against silence and against equally arousing music by another composer on the Stanford-Binet paper folding and cutting task, is powered for a difference of about d = 0.3, and finds a Mozart-specific gap near the claimed 3.56 SAS points. That would move the verdict to OPEN. Nothing located in the 2023 to 2026 record is such a study.

    We could show that the overall estimated effect is small in size (d = 0.37, 95% CI [0.23, 0.52]) for samples exposed to the Mozart sonata KV 448 and samples that had been exposed to a non-musical stimulus or no stimulus at all preceding spatial task performance. Additionally, calculation of effect sizes for samples exposed to any other musical stimulus and samples exposed to a non-musical stimulus or no stimulus at all yielded effects similar in strength (d = 0.38, 95% CI [0.13, 0.63]), whereas there was a negligible effect between the two music conditions (d = 0.15, 95% CI [0.02, 0.28]). Furthermore, formal tests yielded evidence for confounding publication bias, requiring downward correction of effects. … On the whole, there is little evidence left for a specific, performance-enhancing Mozart effect.

    J. Pietschnig, M. Voracek and A. K. Formann, “Mozart effect–Shmozart effect: A meta-analysis”, Intelligence 38, 314–323 (2010), abstract, as given in ERIC record EJ882611

    This phenomenon was received with considerable skepticism in the scientific community and ultimately demonstrated to be a consequence of low study power and bias-related measurement artifacts

    S. Oberleiter and J. Pietschnig, Scientific Reports 13, 3175 (2023), introduction, CC BY 4.0. The paper itself is a meta-analysis about epilepsy; it is cited here only for this dated sentence about the 1993 spatial claim.

    The claimants’ own later position gives both halves. In 2006 Rauscher and Sean C. Hinton wrote that the 1993 finding “was that one specific composition of Mozart enhanced adult spatial test performance for up to about 15 min. There was no indication that other Mozart pieces would have this effect or that the effect was in any way specific to Mozart.” They cited a 2000 meta-analysis by Lois Hetland of “36 studies involving 2,465 subjects” that found the effect “moderate and robust” but limited to one kind of spatial task; and they wrote that “Other studies also support the conclusion that the Mozart effect is largely due to arousal or mood rather than to Mozart or the specific composition”, and that “we agree with Waterhouse that educational practice should not be influenced by this area of research.” (Educational Psychologist 41, 233–238.) The same paper names the popular misreading: a “Mozart effect industry” and a governor’s proposal to send every newborn home with a classical music CD. The 1993 letter claimed nothing about babies or general intelligence.

    The page’s own result, in one sentence: Planted at its printed size of 3.56 Stanford-Binet standard-score points, the 1993 Mozart gap would have been found by the 1999 Western Ontario control (its recomputed F rises from 2.05 to 26.2), though not by the Montreal arm alone; as measured, the two Stanford-Binet arms of that control put the gap at 0.49 of those points, with a standard error of 0.79.

    The dated record. The filled dot marks the decision the verdict dates; the others are milestones.

    1. Rauscher, Shaw and Ky, Nature 365, 611: 36 students, three conditions, the 8 to 9 point translation.
    2. Steele, Bass and Crook, Psychological Science 10, 366-369: 125 participants, Mozart against silence and Philip Glass, treatment F(2, 122) = 0.11, p = .89 (as printed; raw item counts, so it stays out of this page’s power statement).
    3. Nature prints Chabris’s meta-analysis, Steele and colleagues’ three-laboratory replication, and Rauscher’s reply, under one title: “Prelude or requiem for the ‘Mozart effect’?” The replication is the control this page tests; it was answered in the same issue, so the page dates it as a milestone, not the decision.
    4. Hetland, Journal of Aesthetic Education 34 (3/4), from p. 105: a supportive synthesis, described by Rauscher and Hinton in 2006 as 36 studies and 2,465 subjects finding the effect moderate and robust but limited to one kind of spatial task.
    5. Rauscher and Hinton, Educational Psychologist 41, 233-238: the effect is “largely due to arousal or mood rather than to Mozart”, and educational practice should not be influenced by it.
    6. Pietschnig, Voracek and Formann, Intelligence 38, 314-323: d = 0.37 for Mozart against no music, 0.38 for other music, 0.15 between the two, evidence of publication bias, and, as its “central finding”, “the noticeably higher overall effect in studies performed by Rauscher and colleagues than in studies performed by other researchers”. This is the decision the verdict dates.
    7. Oberleiter and Pietschnig, Scientific Reports 13, 3175, describe the 1993 spatial phenomenon as “ultimately demonstrated to be a consequence of low study power and bias-related measurement artifacts”.
    8. Meunier, Tirelli, Durieux and Léonard, preliminary poster: every meta-analytic model gives a small effect, Hedges’ g from 0.099 to 0.236.
    9. Guerrero-Tánori and co-authors, MethodsX 16, 103964: a new EEG protocol with 12 people per group, data on request. A protocol, not a decisive replication.

    V · the further result

    Where the effect lives

    The 2010 meta-analysis did not stop at its average. Its abstract calls its “central finding” something else: “the noticeably higher overall effect in studies performed by Rauscher and colleagues than in studies performed by other researchers, indicating systematically moderating effects of lab affiliation.” The 38 studies it pooled for Mozart against no music are published, with that moderator coded, in the R package metaviz 0.3.1 (GPL-2), whose documentation calls them “A dataset consisting of 38 empirical studies used in the meta-analysis of Pietschnig, Voracek, and Formann (2010)”. Martin Voracek is an author of both. It is a secondary compilation of effect sizes, not the claimants’ data.

    Pooled by random effects (DerSimonian and Laird), the 38 rows give d = 0.372, 95% interval 0.225 to 0.520, which rounds to the printed 0.37 [0.23, 0.52]; restricted maximum likelihood gives 0.375 (0.220 to 0.530). The compilation codes the 1993 study as d = 1.50 (s.e. 0.44); the page could not reproduce that value from the printed 1993 numbers and does not use it as the claimed size. The moderator is coded as “Was the study conducted in the lab of authors Rauscher or Rideout?”. Of the 10 studies coded yes, 5 carry Rauscher’s name, 4 carry the name of B. E. Rideout, an early replicator whose positive results Rauscher’s 1999 reply counts among the replications, and one, Cooper (2004), carries neither. The 2010 abstract speaks of “studies performed by Rauscher and colleagues”. So the split below is the compilation’s coding, and the page calls its two groups by that coding. Split by it, the Rauscher and Rideout studies give 0.880 (0.539 to 1.221), and the other 28 give 0.235 (0.096 to 0.374): a gap of 0.645, the page’s recomputation of the 2010 paper’s finding.

    The 38 studies behind the 2010 meta-analysis

    Studies pooled
    Estimator
    The compilation’s row “Steele, Dalla Bella, et al (1999)”, read as Appalachian State

    d = 0.372, 95% interval 0.225 to 0.520, 38 studies (DL)
    Rauscher and Rideout laboratories 0.880, others 0.235: a gap of 0.645
    1 of 10,000 random relabellings reach a gap of 0.645 or more (seed 448)

    Forest plot of the 38 studiesOne square and 95% interval per study, coloured by whether the compilation codes it as run in Rauscher’s or Rideout’s laboratory, with pooled diamonds for all studies, the Rauscher and Rideout laboratories and the others.
    coded as run in Rauscher’s or Rideout’s laboratoryother laboratoriesa row recoded by you
    The 38 rows as a table
    studynds.e.Rauscher or Rideout labunpublished
    Drawn from the frozen file when the page’s scripts run.

    The moderator on nothing

    A gap between two groups of studies can arise from how the groups were drawn. So the page gives the compilation’s laboratory label to 10 studies chosen at random, reruns the same pooling, and repeats that 10,000 times with a fixed seed (448). A gap of 0.645 or more came up 1 time in 10,000 (an exact one-sided 95% upper bound of 0.00047 on the rate; the largest random gap was 0.735). The difference between laboratories is very unlikely to be a product of which studies happened to be grouped. It does lean on some studies more than others: leave any one of the 38 out and the gap runs from 0.441 (without Rauscher and Ribar (1999) (1), which alone moves it by 0.204) to 0.727 (without Rauscher and Hayes (1999)). Even at its smallest it arises in only 26 of 10,000 relabellings of the remaining studies (upper bound 0.0036). The relabelling cannot say why the groups differ.

    One reading: the protocol

    The Rauscher and Rideout laboratories may run the procedure as designed where others do not. That is the claimants’ 1999 argument: that the replications changed the design and the controls. A 2016 paper in the philosophy of science (Uljana Feest, Studies in History and Philosophy of Science 58, 34–45) uses this very case to discuss the tacit knowledge an experiment can depend on.

    The other reading: the laboratory

    Effects reported from these laboratories may run larger for reasons unrelated to the music. That is the 2010 paper’s “lab affiliation” reading, set beside its own formal evidence of publication bias. It says nothing about intent, and neither does this page.

    What would separate the two readings is the study the verdict asks for: a blind, preregistered, multi-laboratory replication that includes the originating laboratory, with every laboratory running the same protocol.

    Two rows the page could not reproduce

    The compilation’s row “Steele, Dalla Bella, et al (1999)” has n = 86 and d = 0.855 (s.e. 0.23). Its n matches the Appalachian State arm’s total over all five groups (86), the way the compilation’s Western Ontario row counts all three of that laboratory’s groups (67), and the row’s name points to the Nature letter. But it also matches the Mozart and silence groups of Steele, Bass and Crook’s July 1999 replication (44 and 42), which Chabris’s table lists with N = 86 and d = 0.057, and which has no row under its own name in the compilation. The page tests the Appalachian State reading. Neither reading reproduces 0.855: beside Chabris’s 0.057 for Steele, Bass and Crook, Appalachian State’s posttest Mozart minus silence is d = −0.186 (s.e. 0.339), and Mozart against the mean of its four other groups is about −0.08. The coders may have had information the page does not. Replaced by the posttest value, the all-study estimate moves from 0.372 to 0.343 and the other-laboratory estimate from 0.235 to 0.197; dropped, to 0.356 and 0.208. The Rauscher and Rideout estimate does not move. The other questioned row is Chabris’s Western Ontario row in Act II, where recoding moves the critics’ number toward the claim. The compilation’s Montreal row (d = −0.408, n = 32) is close to the page’s −0.403; its Western Ontario row (d = 0.494, n = 67) counts all three of that laboratory’s groups, which reads as Mozart set against both of the others; the page gets 0.518 for that contrast, close but not equal.

    The page adds no publication-bias test of its own; the 2010 paper’s tests and their inputs are not reproduced here.

    The further result, and how far it goes

    The citable results are two. First, planted at its printed size in its own units, the 1993 gap would have been found by the 1999 Western Ontario control (its F from 2.05 to 26.2; power 0.96 for the gap itself) and by the pooled SAS arms (power 0.99), but not by Montreal alone (power 0.61). Second, the 2010 lab-affiliation gap of 0.645 arises in 1 of 10,000 random relabellings. On what exists already, we searched the web through a general search engine (four query forms), PubMed, ERIC, the Figshare index, the CRAN package list and this site’s own index on 2026-09-23 and did not find an interactive page or published analysis that plants the 1993 gap, in its own Stanford-Binet units, into the 1999 control’s printed tables and reports which laboratory could have seen it, or one that reruns the 2010 lab-affiliation split on randomly assigned labels with the coding of each questioned row open to the reader. Prior work the page builds on and credits: the metaviz package itself and a 2024 PLOS figure that plots part of the compilation (static); Fudin and Lembessis (Perceptual and Motor Skills 98, 389–405, 2004), who raised questions about the 1993 scoring of the Stanford-Binet, the experimental design, the validity of its IQ measure and the statistical analyses (abstract read, paper not); Meunier and colleagues’ 2025 poster, which estimates the power of the primary studies in this literature at the meta-analytic effect sizes (not at the 1993 gap, and not by planting it into a control); and a 2006 thesis by R. M. Sweeny (University of Notre Dame) that reran the experiment three times and explored a Bayesian reading of null results.


    VI · the check

    The check

    Recomputed in your browser, now

    The live check runs when the page’s scripts load.

    What this page rests on, and what it chose