Gamified learning for educators & corporate teams

Creativity training works. Just not as well as anyone told you.

Forty activities you can run tomorrow — each one carrying its actual effect size, its facilitator script, and the boundary condition that makes it fail. Including the ones the evidence does not support.

Overhead view of a workshop table covered in blank index cards at the end of a session
Where creativity training actually sits Standardised mean difference (Cohen's d / Hedges' g)
00 — About this resource

I'm not reinventing the wheel. I'm showing you where it came from.

This is a teaching resource, built for educational and non-commercial use. It contains no original research. Every finding here belongs to the researchers credited on each activity — my contribution is curation, evidence grading, and translation into something a facilitator can actually run on a Tuesday morning.

The gap I am trying to close is not a gap in knowledge. The creativity literature has been productive for seventy years. The gap is in transmission: the same handful of unevidenced exercises circulate through classrooms and corporate workshops year after year, while the meta-analyses that would settle the question go unread. So this site does three things — it names the primary study behind each activity, it grades how much weight that study can bear, and it says plainly when a popular technique does not have the evidence people assume.

If you use any of it, cite the original researchers rather than this site. Full citations sit inside each activity card, and a complete reference list is available in the companion guide. Free under CC BY-NC-SA — take it, adapt it for your own teaching, and correct me where I have got something wrong.

On checking my working. Every activity carries a resolver link to its primary source and, where one exists, a DOI. The evidence grade on each card is not a score I have assigned by feel — it is three separately stated judgements about design, replication and outcome, defined in full below the filters. If a link does not resolve, or you think a grade is wrong, that is a correctable error and I would like to hear about it.

01 — Read this first

The number the field spent fifty years getting wrong

Before teaching any creativity activity, you need to know what the field's own best evidence says about creativity training as a whole. This is what separates an evidence-informed programme from a Pinterest board.

Sio & Lortie-Forgues (2024), in Psychological Bulletin, is the largest meta-analysis of creativity training ever conducted: 169 studies, 844 effect sizes, 18,486 participants, five decades. Critically, it included 48 unpublished studies specifically to counteract publication bias.

0.53Raw effect size — in line with the five earlier meta-analyses (0.47–1.02)
0.29–0.32After four independent bias corrections. A 39–45% reduction
11%Of 169 studies that used randomisation, an active control and a pretest
1Preregistered study. In fifty years. Nine were replications, seven predate 2010
Three defensible claims come out of this. First: creativity training works, but modestly — and unlike ego depletion or growth mindset, the corrected effect stayed significantly above zero. An effect near 0.3 beats the mean educational intervention, which Kraft's (2020) review of 747 randomised trials put at d = 0.16 — though the median in that review was 0.10, and the mean is the more flattering of the two numbers to quote. Second: cognitive-skills training beat motivational training, and this survived the full multi-moderator model. Teach people how to think, not how to feel creative. Third: anyone promising a 300% creativity boost is selling something.
The number this site has to answer for. Look at the leftmost mark on the ruler above. Restricted to the nineteen studies that met the highest design standard, and corrected for bias, the estimate falls to roughly d = 0.11 — below the typical education intervention on the same axis. That is the uncomfortable reading, and the honest summary is this: the corrected whole-field effect of about 0.3 is real, and the effect shrinks as design quality rises. Both statements are true, they pull in opposite directions, and which one you believe depends on whether you think the weaker studies are adding signal or noise. On the current evidence, the defensible position is that creativity training does something, that the something is smaller than the field advertises, and that nobody has yet run the trial that would settle it. Section 07 describes that trial.
Did you know When a research field consistently reports effects larger than almost everything else in social science, the most likely explanation is not that the field is unusually effective. Lovakov & Agadullina (2021) extracted 6,447 Cohen's d values from published social-psychology meta-analyses: the 25th, 50th and 75th percentiles were 0.15, 0.36 and 0.65. doi:10.1002/ejsp.2752
02 — The library

Forty activities, graded

Filter by evidence strength, by which phase of the process it serves, or by how long you have. Open any card for the run sheet, the facilitator script with timers, and the study behind it. Seven of the forty are on the shelf because the evidence does not support what people believe about them.

How the grading works — three judgements, stated separately, no composite score

The tier badge (Strong / Moderate / Contested) is a summary. The three-part meter under each blurb is the actual grade, and it is deliberately not collapsed into one number, because a single number would hide exactly the distinctions this site exists to make. A framework can be worth teaching and still score zero on outcome; that is information, not a verdict.

D — Design

  1. Qualitative, case-based or theoretical only
  2. Correlational or field-observational
  3. Controlled experiment, single laboratory
  4. Meta-analysis, or trial with an active control

R — Replication

  1. Failed a preregistered replication
  2. Single study, not independently replicated
  3. Independently replicated at least once
  4. Meta-analytic or consistently replicated

O — Outcome

  1. Creativity was never measured
  2. An adjacent cognitive outcome only
  3. A validated creativity measure (AUT, RAT, insight)
  4. Real creative products or field innovation outcomes

Each list runs 0 at the top to 3 at the bottom. Every card states which level it sits at on each dimension and why, under Evidence profile in its Evidence tab.

Evidence Phase
01Strong

Quantity-first ideation

Enforce a ten-minute minimum on idea generation, then show people where their best ideas actually landed.

DRO
25 minAnydiverge
Full detail

How to run it

Give a divergent prompt — a real design problem, or the classic Alternate Uses Task ("list uses for a brick"). Set a ten-minute minimum and enforce it. No stopping at "I'm done." No evaluation, no editing, quantity only.

Afterwards, have each person circle their three most original ideas and note where in the list they fall. Show of hands: how many circled something from the back half?

Facilitator script

  1. 2 min · Set up
    For the next ten minutes there is exactly one rule: quantity. No editing, no judging, no crossing out. If an idea is stupid, write it down anyway — stupid ideas are load-bearing.
  2. 10 min · Silent generation
    Ten minutes on the clock. When you feel like you have run out, that is the halfway point, not the end. Keep going.
  3. 3 min · Self-selection
    Now go back through your list and circle the three you think are most original. Not most useful — most original.
  4. 8 min · The reveal
    Hands up if any of your three came from the first half of your list. Now the second half. Look around the room. That gap is one of the oldest findings in creativity research, and you just reproduced it in ten minutes.
  5. 2 min · Land it
    The implication is uncomfortable. Every brainstorm you have ever cut short at the natural energy dip ended precisely before the good ideas arrived.

The evidence

The serial order effect is one of the oldest and most robust findings in creativity research, dating to Christensen, Guilford & Wilson (1957). Beaty & Silvia (2012) time-stamped every response from 133 adults across a ten-minute unusual-uses task and had three raters score each for creativity. Originality rises across time while fluency declines. It has replicated in adults, older children, and five-to-six-year-olds.

Early responses come from direct memory retrieval — known uses, seen uses. Only once that well runs dry do people engage the effortful executive processes that produce genuinely remote associations. The EEG picture matches: alpha power shows a U-shaped curve, dipping mid-task before rising again ahead of the more original responses.

Evidence profile

DesignMeta-analysis, or trial with an active control
ReplicationMeta-analytic or consistently replicated
OutcomeA validated creativity measure (AUT, RAT, insight)

Cognitive-skills ideation training is the moderator the 169-study meta-analysis supports; outcome is divergent-thinking scoring, not real output.

Beaty, R. E., & Silvia, P. J. (2012). Why do ideas get more creative across time? Psychology of Aesthetics, Creativity, and the Arts, 6(4), 309–319. doi:10.1037/bul0000432 Find this paper ↗

Did you knowAlmost everyone circles ideas from the back half. This is the single most reliable "aha" demonstration in creativity education, because participants discover the finding in their own handwriting rather than being told it.
Watch outQuantity-first assumes Osborn's defer-judgment rule, and that rule is contested. Nemeth et al. (2004) found instructed debate outperforming it — see activity 27, which this library carries precisely so you are not choosing between them blind. The current best reconciliation is that the damage in group brainstorming comes from production blocking and evaluation apprehension rather than from disagreement as such; that reading is defensible, not settled.
02Moderate

Deliberate constraint injection

Split the room. Half get the open brief, half get the same brief plus an arbitrary restriction. Rate blind.

DRO
35 min8–40diverge
Full detail

How to run it

Split the group. Half receive the open brief ("design a better commute"). Half receive the same brief plus an arbitrary constraint — "and it must involve a bucket", "using only things already in this building", "in under fifty words".

Collect outputs, strip the condition labels, and have a third party rate them for creativity. Reveal the split afterwards.

Facilitator script

  1. 3 min · Distribute
    You each have a brief. Do not compare them yet. Some of you have an extra line — that is deliberate, and I will explain at the end.
  2. 12 min · Generate
    Twelve minutes. Work individually. Whatever is on your sheet is the whole task.
  3. 8 min · Blind rating
    Pass your sheet to the far side of the room. Score each one you receive from one to five on originality only. Do not score usefulness.
  4. 8 min · Reveal and discuss
    Now compare briefs with the person next to you. Which group scored higher? Almost every time, it is the group who had less freedom.
  5. 4 min · Transfer
    So the practical question is not how do we remove obstacles. It is which constraint do we choose, and do we choose it deliberately or let the org chart choose it for us.

The evidence

Haught-Tromp (2017), Psychology of Aesthetics, Creativity, and the Arts: participants wrote two-line greeting-card rhymes. On half the trials they were additionally required to include a specific, arbitrary noun. Constrained rhymes were rated more creative, in both studies. An order interaction also indicated a carryover effect — practice with constraints improved performance on later unconstrained tasks.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationIndependently replicated at least once
OutcomeA validated creativity measure (AUT, RAT, insight)

Controlled experiments plus a meta-analytic review of the inverted-U; outcomes are rated idea originality.

Haught-Tromp, C. (2017). The Green Eggs and Ham hypothesis: How constraints facilitate creativity. Psychology of Aesthetics, Creativity, and the Arts, 11(1), 10–17. Find this paper ↗

Did you knowThe study is named after Green Eggs and Ham. Theodor Geisel's publisher bet him he could not write a children's book using fifty distinct words or fewer. Geisel won the bet and produced one of the best-selling children's books ever written. The constraint did not limit the work — it was the work.
Watch outTask constraints help. Social constraints — surveillance, expected evaluation, contingent reward — reliably hurt. Constrain the problem, never the person.
03Strong

Far-analogy transfer

Strip the problem of all domain language, then raid distant fields for mechanisms.

DRO
45 min6–24diverge
Full detail

How to run it

State the problem as an abstract function, stripping every piece of domain vocabulary. "How do we onboard new employees faster" becomes "how does a system rapidly integrate a foreign element without rejecting it."

Then ask who else solves that — immunology, coral reefs, jazz ensembles, aviation, prison systems, beekeeping. Harvest three mechanisms from each distant domain and translate them back.

Facilitator script

  1. 8 min · Strip the language
    Write your problem on the board. Now remove every noun specific to your industry. If a stranger could not tell what business you are in from the sentence, you have done it right.
  2. 5 min · Draw domains
    Each table draws three domain cards. You do not get to swap. The further from your world, the better this works.
  3. 15 min · Mine mechanisms
    How does your domain solve the abstract problem? I want mechanisms, not vibes. Not "bees are collaborative" — how does a hive actually allocate foragers to a food source?
  4. 12 min · Translate back
    Now map it. What is the equivalent of the waggle dance in your onboarding process? Be literal. The literalness is where the idea is.
  5. 5 min · Debrief
    Notice which analogies produced something and which produced decoration. The difference is always whether you mapped function or surface.

The evidence

Jeppesen & Lakhani (2010), Organization Science, analysed 166 R&D challenges broadcast to over 12,000 scientists on InnoCentive. The probability of submitting the winning solution rose with increasing distance between the solver's field of expertise and the problem's field. Outsiders won not despite their distance but because of it — they imported perspectives and heuristics unavailable inside the field.

Evidence profile

DesignMeta-analysis, or trial with an active control
ReplicationMeta-analytic or consistently replicated
OutcomeA validated creativity measure (AUT, RAT, insight)

Multiply replicated in analogical-transfer experiments; the field-outcome evidence sits in activity 09 rather than here.

Jeppesen, L. B., & Lakhani, K. R. (2010). Marginality and problem-solving effectiveness in broadcast search. Organization Science, 21(5), 1016–1033. Find this paper ↗

Did you knowThe same study found female solvers — described by the authors as being in the "outer circle" of the scientific establishment — significantly outperformed men at producing winning solutions. Social marginality functioned the same way technical marginality did.
Watch outDistance only pays when the analogy is mapped at the level of function, not surface features. "Our app should be like a beehive because both are busy" is decoration. "Our app should use waggle-dance-style peer signalling to allocate attention" is transfer.
04Strong

Structured incubation

Load the problem properly, then break with a low-demand task. Not rest. Not email. Washing up beats staring at a wall.

DRO
55 minAnydiverge
Full detail

How to run it

Front-load the work: 20–30 minutes of genuine, effortful attempt at the problem. Then break for 15–20 minutes with a low-cognitive-demand task — a walk, tidying, undemanding admin. No email, no meetings, no problem-adjacent thinking. Return and re-attempt.

Facilitator script

  1. 25 min · Load the problem
    Twenty-five minutes of real attempt. Not planning to attempt — attempting. You need to hit a wall for the next part to work at all.
  2. 18 min · Low-demand break
    Now put it down completely. Go for a walk, tidy something, do the dishes. Not your phone. Not your inbox. The task needs to occupy your hands and almost none of your mind.
  3. 12 min · Re-attempt cold
    Straight back in, no warm-up, no recap. Write for two minutes before you speak to anyone.

The evidence

Sio & Ormerod (2009), Psychological Bulletin, meta-analysed 117 studies. Overall incubation effect: d = 0.29. Three moderators matter for how you run it. Divergent thinking tasks benefit most, more than linguistic or visual insight problems. A low-demand break task beats a high-demand task and beats resting with no task at all. And longer preparation before the break produces a larger effect.

Evidence profile

DesignMeta-analysis, or trial with an active control
ReplicationMeta-analytic or consistently replicated
OutcomeA validated creativity measure (AUT, RAT, insight)

Meta-analysis of incubation studies with moderator analysis; outcomes are lab divergent and insight tasks.

Sio, U. N., & Ormerod, T. C. (2009). Does incubation enhance problem solving? A meta-analytic review. Psychological Bulletin, 135(1), 94–120. doi:10.1037/a0014212 Find this paper ↗

Did you knowThe finding that a mildly occupied mind beats a completely idle one is the practically useful part. "Go stare at a wall" underperforms "go do the washing up."
Watch outTeams routinely schedule the break before doing the preparation. Without sufficient prior engagement there is nothing to incubate — and incubation becomes an expensive word for procrastination.
05Strong

Walking ideation

Move the divergent phase onto its feet. Then sit down for the convergent phase — walking makes that slightly worse.

DRO
20 minPairsstate
Full detail

How to run it

Move divergent thinking sessions onto their feet. Either a corridor loop or an outdoor walk. Pairs work well: one walks and talks, the other captures. If you cannot walk during the session, walk immediately before it — the effect persists.

Sit down again before any decision-making.

Facilitator script

  1. 2 min · Pair and brief
    Pair up. One of you talks, one of you writes — and the writer does not edit, does not improve, does not filter. You are a recording device with legs.
  2. 8 min · Walk, person A
    Eight minutes. Keep moving the whole time. If the conversation stalls, keep walking through the silence rather than stopping.
  3. 8 min · Walk, person B
    Swap. Same rules.
  4. 2 min · Sit down
    Now sit. Deliberately. We are switching from generating to choosing, and the research says those want different physical states.

The evidence

Oppezzo & Schwartz (2014), Journal of Experimental Psychology: LMC, four experiments, 176 participants. Walking increased divergent thinking on Guilford's Alternate Uses for 81% of participants in Experiment 1; across the four experiments the proportion of walkers who improved was 81%, 88%, 95% and 100%. The effect held indoors on a treadmill facing a blank wall as well as outdoors, so it is not environmental stimulation. There was a residual effect: people who sat down after walking still outperformed people who sat throughout.

Critically, walking did not help convergent thinking — Compound Remote Associates performance was mildly worse.

Evidence profile

DesignMeta-analysis, or trial with an active control
ReplicationIndependently replicated at least once
OutcomeA validated creativity measure (AUT, RAT, insight)

Four within-subject experiments with control conditions, independently extended; outcome is AUT and analogy generation.

Oppezzo, M., & Schwartz, D. L. (2014). Give your ideas some legs. Journal of Experimental Psychology: Learning, Memory, and Cognition, 40(4), 1142–1152. doi:10.1037/a0036577 Find this paper ↗

Did you knowThe convergent-thinking result is the most useful part and almost never gets quoted. Walk for the divergent phase; sit down for the phase where you need one right answer. Using the same physical state for both wastes half the effect.
Watch outThe widely circulated "60% increase" figure summarises one measure in one experiment. Cite the proportion of participants who improved instead — better documented, and more honest.
06Moderate

Sleep-onset capture

The Edison protocol. Fifteen seconds in N1 tripled insight rates — but the window closes the moment you fall properly asleep.

DRO
25 minSolostate
Full detail

How to run it

Best taught as a demonstrated technique participants try at home rather than a workshop activity.

Load the problem deliberately. Sit semi-reclined holding a light object over a hard surface. Doze. When the object falls, immediately write down whatever was in your mind — before you evaluate any of it.

Facilitator script

  1. 5 min · Load
    Spend five minutes deliberately turning the problem over. You are not trying to solve it. You are trying to make it the most recently active thing in your head.
  2. 15 min · Drift
    Semi-recline, object in hand over a hard floor. Let yourself go. The drop will wake you.
  3. 5 min · Capture immediately
    Write before you judge. Hypnagogic material reads as nonsense on the way in and turns out to be a metaphor on the way out. Judge it tomorrow.

The evidence

Lacaux et al. (2021), Science Advances, N = 103. Participants were given maths problems containing a hidden shortcut rule. During a 20-minute rest with EEG monitoring, those who spent as little as 15 seconds in N1 — the hypnagogic sleep-onset stage — discovered the rule at a rate of 83% versus 30% for those who stayed awake. The benefit disappeared if participants slipped past N1 into deeper N2 sleep.

Evidence profile

DesignMeta-analysis, or trial with an active control
ReplicationSingle study, not independently replicated
OutcomeA validated creativity measure (AUT, RAT, insight)

Well-controlled polysomnography study with a clean comparison, but a single sample and a rule-discovery outcome rather than open ideation.

Lacaux, C., Andrillon, T., et al. (2021). Sleep onset is a creative sweet spot. Science Advances, 7(50), eabj5866. doi:10.1126/sciadv.abj5866 Find this paper ↗

Did you knowThe 2021 study tested whether Edison's dropped-object trick actually catches N1, and found it an imperfect but real marker: 81% of participants who dropped the object reported drifting off, and EEG showed a clear delta-power rise in the ~24 seconds before the drop. Of 63 droppers, though, 26 had already passed through N1. Edison was right about the window and roughly right about the method.
Watch outSingle study, laboratory task, one specific kind of insight problem. Teach it as a fascinating, plausible technique — not as settled science.
07Strong

Brainwriting before talking

Never open with a verbal brainstorm. Production blocking is mechanical, not motivational — and it scales with group size.

DRO
40 min4–30diverge
Full detail

How to run it

  1. Silent generation — ten minutes, individually, in writing
  2. Round-robin capture — each person reads one idea at a time, no discussion, until all are on the wall
  3. Clarify — questions of understanding only, no evaluation
  4. Build — now go verbal, using the pooled set as raw material
  5. Converge — see activity 12

Facilitator script

  1. 2 min · Set the rule
    We are not going to brainstorm out loud, and I will tell you why afterwards. For now: silence, and one idea per card.
  2. 10 min · Silent generation
    Ten minutes. One idea per card. Do not talk, do not look at anyone else's cards, do not stop early.
  3. 12 min · Round-robin
    Going round the circle, one card each, read it and put it up. No commentary, no "oh I had that too", no reactions. We are building the pool, not judging it.
  4. 6 min · Clarify only
    Questions of understanding only. You may ask what someone meant. You may not ask whether it would work.
  5. 10 min · Now talk
    Now the group conversation is worth having, because everybody's ideas are already in the room instead of trapped behind whoever spoke first.

The evidence

Diehl & Stroebe (1987) ran four experiments isolating three candidate mechanisms for the productivity loss in brainstorming groups — free riding, evaluation apprehension, and production blocking — and found that production blocking accounted for most of it. Mullen, Johnson & Salas (1991) meta-analysed the literature and confirmed that nominal groups reliably outproduce interacting brainstorming groups, with the gap widening as group size grows.

Production blocking is mechanical. Only one person can speak at a time, so while you wait your turn you either forget your idea, suppress it as no-longer-relevant, or lose your thread listening to someone else. In a six-person group each member spends the large majority of the session listening rather than generating.

Evidence profile

DesignMeta-analysis, or trial with an active control
ReplicationMeta-analytic or consistently replicated
OutcomeA validated creativity measure (AUT, RAT, insight)

Meta-analytic evidence on production blocking and nominal-versus-real groups, replicated for four decades.

Diehl, M., & Stroebe, W. (1987). Productivity loss in brainstorming groups. Journal of Personality and Social Psychology, 53(3), 497–509. doi:10.1037/0022-3514.53.3.497 Find this paper ↗

Did you knowGroups consistently rate themselves as more productive than individuals working alone — the illusion of group productivity — even when the objective count of unique ideas shows the opposite. The subjective experience of a brainstorm is genuinely energising. That is precisely why the practice has survived seventy years of contrary evidence.
Watch outDo not over-correct into "meetings are useless." Groups are excellent at combination, elaboration and selection. They are poor at generation. Sequence accordingly.
08Strong

Problem reframing ladder

Ban solutions for thirty minutes. Generate problem statements instead, then vote on which one to ideate against.

DRO
30 min4–30frame
Full detail

How to run it

Ban solutions for the first thirty minutes. Instead, generate problem statements. Take the presenting problem and rewrite it at five levels of abstraction, then generate at least ten alternative framings. Only then vote on which framing to ideate against.

Facilitator script

  1. 3 min · Impose the ban
    For the next half hour, nobody proposes a solution. If a solution arrives, write it on a card, put it face down, and come back to it later. It will still be there.
  2. 10 min · Climb the ladder
    Take your problem and ask "why does that matter?" five times, writing each answer as a new problem statement. You are climbing towards the abstract.
  3. 10 min · Ten framings
    Now ten alternative framings, anywhere on that ladder. Some should feel too broad and some too narrow. That is the range we want to see.
  4. 7 min · Choose deliberately
    Vote. And notice what you are doing — you are choosing the shape of every idea you will have for the rest of the day.

The evidence

This is the practical expression of the strongest moderator in the Sio & Lortie-Forgues meta-analysis: cognitive-content training outperformed non-cognitive training, and this survived the full multi-moderator model (p = .024). Problem construction is the most cognitive of the creativity skills — it is the part where an actual thinking procedure is taught. It maps onto Mumford's problem-construction research programme, which consistently finds that how a problem is defined constrains the entire downstream solution space.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationIndependently replicated at least once
OutcomeA validated creativity measure (AUT, RAT, insight)

Controlled problem-construction experiments, independently replicated; outcomes are rated solution quality and originality.

Sio, U. N., & Lortie-Forgues, H. (2024). The impact of creativity training on creative performance. Psychological Bulletin, 150(5), 554–585. Find this paper ↗

Did you knowAway-days designed to make people feel creative are the non-cognitive category. Sessions that teach a specific thinking procedure are the cognitive category. The evidence favours the second — and it is usually cheaper, which is a useful thing to be able to say in a budget conversation.
Watch outTeams find this phase intensely uncomfortable. Expect resistance around minute twelve. That discomfort is the fixation being broken; if you rescue them from it, you have wasted the exercise.
09Moderate

Outsider seat

Deliberately seat someone with no domain expertise, and brief them that naive questions are the job.

DRO
15 minAnyclimate
Full detail

How to run it

When assembling an ideation team, deliberately include at least one person with no domain expertise in the problem, and brief them explicitly that naive questions are their job, not a liability. Pair them with a domain expert acting as translator rather than gatekeeper.

Give the outsider a formal right of interruption. Without it, social pressure silences them inside ten minutes.

Facilitator script

  1. 5 min · Brief the outsider privately
    Your job today is to ask the questions everyone else is too embarrassed to ask. If you understand everything being said, you are not doing your job.
  2. 2 min · Announce the role
    Sam is here as our outsider. Sam has a standing right to stop us at any point. When Sam does, we answer properly — no shorthand, no acronyms.
  3. 8 min · Harvest at the end
    Sam, which of our assumptions sounded strangest to you? Not wrong — strange. Those are the ones worth checking.

The evidence

Jeppesen & Lakhani (2010) as above. Also relevant: Maddux & Galinsky (2009), JPSP, found across five studies that time spent living abroad predicted creative performance, that the effect was causal, and that it was mediated by degree of cultural adaptation rather than mere exposure. Maddux, Adam & Galinsky (2010) sharpened this: the active ingredient is functional multicultural learning — understanding why people in another culture do what they do, not just observing that they do it.

Evidence profile

DesignCorrelational or field-observational
ReplicationSingle study, not independently replicated
OutcomeReal creative products or field innovation outcomes

Large field dataset with real winning solutions, but observational — solvers were not randomly assigned a distance from the domain.

Maddux, W. W., & Galinsky, A. D. (2009). Cultural borders and mental barriers. Journal of Personality and Social Psychology, 96(5), 1047–1061. doi:10.1287/orsc.1090.0491 Find this paper ↗

Did you knowPriming people to recall foreign living experiences boosted creativity — but only for people who had actually lived abroad. For everyone else the prime did nothing. You cannot shortcut the experience; there has to be something in memory to prime.
Watch outTourism is not adaptation. Neither is a diversity statement. The mechanism is the effortful business of having your default assumptions fail and rebuilding them.
10Strong

Building psychological safety

The most-cited corporate finding in this space is essentially a replication of a 1999 academic paper.

DRO
60 minIntact teamsclimate
Full detail

How to run it

Standard, well-tested levers: leaders model fallibility by naming their own errors first; frame the work explicitly as a learning problem rather than an execution problem; respond to bad news with curiosity rather than blame; separate idea generation from idea ownership so that critique cannot be read as personal attack.

Facilitator script

  1. 10 min · Leader goes first
    I am going to start by describing a call I got wrong this quarter, and what it cost. Nobody else speaks until I am finished.
  2. 15 min · Reframe the work
    Is this a problem where we already know the answer and need to execute, or one where we do not and need to learn? Be honest. Most teams claim the second and behave like the first.
  3. 20 min · Detach ideas from owners
    From here on, ideas go on the wall unattributed. You may attack any idea on that wall as hard as you like. You may not attack a person.
  4. 15 min · Test it
    Someone name something that is going worse than we have been saying out loud. I will respond with questions, not consequences. Watch what I do — that is the actual policy.

The evidence

Frazier et al. (2017), Personnel Psychology, conducted a meta-analytic review confirming that psychological safety significantly predicts creativity, voice behaviour, information sharing, learning behaviour and task performance, at both individual and team levels. Edmondson (1999) established the construct; Newman, Donohue & Eva (2017) systematically reviewed the literature and mapped its antecedents.

A necessary nuance: recent work on nonlinear relationships finds the benefits concentrated in non-routine, exploratory work. For highly standardised tasks the returns flatten. Psychological safety is a creativity intervention, not a universal performance intervention.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationMeta-analytic or consistently replicated
OutcomeReal creative products or field innovation outcomes

Field studies and meta-analysis with real team outcomes; the design is predominantly correlational rather than experimental.

Frazier, M. L., et al. (2017). Psychological safety: A meta-analytic review and extension. Personnel Psychology, 70(1), 113–165. doi:10.2307/2666999 Find this paper ↗

Did you knowGoogle's Project Aristotle studied hundreds of its own teams looking for the composition variables that predicted performance — the right mix of skills, seniority, personality. Composition explained very little. The strongest predictor was psychological safety.
Watch outYou cannot facilitate this into existence in a workshop. What you can do in a workshop is demonstrate the levers and show the leader what their own behaviour is currently teaching the room.
11Strong

Naming the bias against creativity

Under uncertainty, people develop an implicit bias that stops them recognising a creative idea — while sincerely wanting one.

DRO
20 minDecision-makersconverge
Full detail

How to run it

A twenty-minute teaching module, ideally delivered to decision-makers rather than ideators. Present the finding, then audit your own organisation's last five rejected proposals against it.

Then introduce structural counters: blind evaluation, mandated advocacy for the most novel option, and separating the uncertainty-reduction conversation from the novelty-assessment conversation.

Facilitator script

  1. 5 min · Present the finding
    People reject novel, high-quality ideas while sincerely reporting that they want creative ideas. Not hypocrisy — an implicit bias that switches on under uncertainty.
  2. 10 min · Audit your own
    Pull up the last five proposals this group turned down. For each one, ask: was the stated reason about the idea, or about our uncertainty? Be uncomfortable about this.
  3. 5 min · Install a counter
    From today, every shortlist carries one high-novelty option with a named advocate whose job is to argue for it. Not to believe in it — to argue for it.

The evidence

Mueller, Melwani & Goncalo (2012), Psychological Science. Two experiments manipulating uncertainty, including an uncertainty-reduction prime. Results showed a negative implicit bias against creativity relative to practicality whenever participants experienced uncertainty — and, more damagingly, this bias impaired participants' ability to recognise a creative idea when they saw one.

Evidence profile

DesignMeta-analysis, or trial with an active control
ReplicationIndependently replicated at least once
OutcomeA validated creativity measure (AUT, RAT, insight)

Randomised experiments with implicit measures, independently replicated; outcome is evaluation of creative ideas, not production.

Mueller, J. S., Melwani, S., & Goncalo, J. A. (2012). The bias against creativity. Psychological Science, 23(1), 13–17. doi:10.1177/0956797611421018 Find this paper ↗

Did you knowThe authors describe the irony directly. Uncertainty is exactly what drives organisations to seek creative solutions — and it is also exactly the condition that makes them unable to recognise one. Your innovation programme is most likely to be sabotaged at the precise moment it is most needed.
Watch outThis is the highest-leverage module in most corporate innovation curricula and the most commonly omitted. Organisations spend heavily teaching people to generate ideas and almost nothing teaching people to receive them.
12Moderate

Structured selection

Score novelty and feasibility separately, by different people. Combining them into one score is where novel ideas die.

DRO
45 minAnyconverge
Full detail

How to run it

  1. Cluster the raw pool by underlying mechanism, not surface theme
  2. Score on two independent axes — novelty and feasibility — separately, and by different people
  3. Force-select from the novelty quadrant. At least one advancing concept must come from high-novelty / low-feasibility, with a named advocate
  4. Prototype to learn, not to prove. The first prototype converts an uncertainty into a cheap question
  5. Record the rejected set. It is the most valuable artefact of the session and is almost always binned

Facilitator script

  1. 12 min · Cluster by mechanism
    Group these by how they work, not by what they are about. Two ideas in different markets that use the same mechanism belong together.
  2. 10 min · Score novelty only
    This half of the room scores novelty. Only novelty. I do not want to hear the word "realistic".
  3. 10 min · Score feasibility only
    This half scores feasibility. Only feasibility. I do not want to hear the word "exciting".
  4. 8 min · Plot and force-select
    Plot them. Now — one advancing concept has to come from the top-left quadrant. Who is advocating for it?
  5. 5 min · Bank the rejects
    Photograph the whole wall including everything we killed. In six months the constraint that killed half of these will have changed.

The evidence

Divergent thinking is only half of creativity — the definition used across this literature is ideas that are novel and useful. Yet most creativity activities, and most creativity research, address only novelty: in the Sio & Lortie-Forgues sample, 137 of 169 studies used divergent thinking as the outcome and only 29 measured an actual creative product.

Research on idea selection consistently finds that groups choose ideas that are feasible and familiar over ideas that are original — even when explicitly instructed to select the most creative option, and even when they generated the original ideas themselves.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationSingle study, not independently replicated
OutcomeA validated creativity measure (AUT, RAT, insight)

Controlled selection experiments; the specific protocol here is assembled from several studies rather than tested as a whole.

Rietzschel, E. F., Nijstad, B. A., & Stroebe, W. (2010). The selection of creative ideas after individual idea generation. British Journal of Psychology, 101(1), 47–68. Find this paper ↗

Did you knowRejected ideas are the most valuable artefact of an ideation session and are almost always thrown away. The constraint that killed an idea is usually more temporary than the idea itself.
Watch outIf novelty and feasibility are scored by the same person at the same moment, you have not run this activity. You have run a feasibility review with extra steps.
13Strong

Generate / evaluate separation

A ten-minute teaching module with a neural justification: the two modes run on networks that are anticorrelated at rest.

DRO
10 minAnyframe
Full detail

How to run it

Teach the room to name which mode they are in, out loud, before speaking. Give each person two pens — one colour for generative contributions, one for evaluative. Everything written must use the matching colour.

Crude, slightly silly, and extremely effective at making the switching visible.

Facilitator script

  1. 3 min · Explain the architecture
    Generating and evaluating run on brain networks that are anticorrelated at rest. When I ask you to "be creative but realistic", I am asking you to run two opposed systems at once. That is not rigour. That is a design fault in the instruction.
  2. 2 min · Assign the pens
    Blue is generative. Red is evaluative. You may switch as often as you like — but you have to pick up the other pen to do it.
  3. 5 min · Run and observe
    Look at the balance of colour on your sheet. Most teams discover they are running eighty percent red and calling it ideation.

The evidence

Beaty et al. (2018), PNAS: connectome-based predictive modelling on 163 participants identified a network spanning default, salience and executive systems — networks that typically work in opposition. Connectivity strength within it predicted idea originality across four independent datasets.

A 2025 multi-centre study in Communications Biology (10 samples, five countries, N = 2,433) found that the number of dynamic switches between default and executive networks predicted creativity but not general intelligence.

Evidence profile

DesignMeta-analysis, or trial with an active control
ReplicationMeta-analytic or consistently replicated
OutcomeA validated creativity measure (AUT, RAT, insight)

Converging evidence from group ideation experiments and the network-switching neuroimaging literature.

Beaty, R. E., Kenett, Y. N., et al. (2018). Robust prediction of individual creative ability from brain functional connectivity. PNAS, 115(5), 1087–1092. Find this paper ↗

Did you knowThe 2025 finding is the interesting one for educators: switching predicted creativity but not intelligence. Whatever this capacity is, it is not simply being clever — which is the strongest neural argument yet that it can be trained rather than only selected for.
Watch outNeuroscience explains; it does not prescribe. The pens work because they make the switching visible, not because they change anyone's connectivity in a ninety-minute session.
14Moderate

Low-stimulation generation

Idea generation is an inward task with an EEG signature. Most workshop rooms are optimised against it.

DRO
15 minAnystate
Full detail

How to run it

For the generation phase only: dim the lights, close laptops, stop talking, and offer explicit permission to close eyes. No music, no ambient video, no facilitator narration over the silence.

Then bring the light and energy back for the sharing phase. The contrast is the point.

Facilitator script

  1. 2 min · Change the room
    Laptops closed, lights down, phones face down and out of reach. If you want to close your eyes for this, close them.
  2. 8 min · Generate inward
    I am going to stop talking now. I will not fill the silence, and neither should you. Eight minutes.
  3. 5 min · Bring it back up
    Lights up. Now we go loud. Notice how different that felt from a normal workshop — and that the difference was mostly the absence of me talking.

The evidence

EEG alpha power (8–13 Hz) rises reliably during creative idea generation and rises more in more creative individuals; Fink & Benedek's review describes it as among the most consistent findings in creativity neuroscience. Benedek et al. showed alpha increases specifically when bottom-up processing is prevented, supporting the interpretation of alpha as internally-directed attention — the brain damping external input to work with internal material.

That was correlational until Ritter, Abbing & van Schie (2018) tested it directly. Participants completed creativity and working-memory tasks twice, once with eyes open and once with eyes closed. Eye-closure improved performance on both the divergent task (Adapted Alternative Uses) and the convergent task (Remote Associates) — with no effect on working memory. The absence of a working-memory effect is the important part: it indicates the benefit is specific to creativity rather than a general cognitive boost.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationIndependently replicated at least once
OutcomeA validated creativity measure (AUT, RAT, insight)

EEG correlational evidence plus a controlled eye-closure experiment; see activity 38, which points the other way.

Ritter, S. M., Abbing, J., & van Schie, H. T. (2018). Eye-closure enhances creative performance on divergent and convergent creativity tasks. Frontiers in Psychology, 9, 1315. See also Fink, A., & Benedek, M. (2014), Neuroscience & Biobehavioral Reviews, 44, 111–123. Find this paper ↗

Did you knowAlpha follows a U-shaped curve across a generation task: high at the start, dipping in the middle, rising again just before the more original responses appear — most sharply in the most creative participants. The mid-task dip is the felt experience of running dry, and it is a stage rather than a stopping point.
Watch outThis card points in the opposite direction to a well-known published finding: moderate ambient noise around 70 dB improving creative cognition (Mehta, Zhu & Cheema, 2012). That result has a shaky replication record — see activity 38 — but it is not refuted, and this activity is less secure than its grade suggests. Do not present low-stimulation conditions as settled. And never make eye-closure compulsory; for some participants it is genuinely unpleasant.
15Moderate

Schedule against the chronotype

Insight problems are solved better at your NON-optimal time of day. Analytic problems are unaffected.

DRO
5 minPlanning decisionstate
Full detail

How to run it

A planning rule rather than an activity. Poll the group's chronotypes in advance, then schedule the divergent session against the grain — late afternoon for a morning-heavy team, early morning for an evening-heavy one.

Keep the sharp hours for analytic work and decisions, which time of day does not reliably help anyway.

Facilitator script

  1. 2 min · Poll in advance
    Before the session: are you sharpest first thing, or later in the day? One word answers, in the invite.
  2. 3 min · Explain on the day
    You may have noticed I have scheduled our idea generation at four in the afternoon, which feels like a strange choice. It is deliberate, and the reason is that slightly foggy is better for insight than razor sharp.

The evidence

Wieth & Zacks (2011), Thinking & Reasoning: participants solved insight and analytic problems at their optimal or non-optimal time of day. Insight problem-solving was consistently better at the non-optimal time. Analytic problem-solving showed no consistent time-of-day effect. The proposed mechanism is reduced inhibitory control — the same loosening that lets irrelevant material into working memory also lets remote associations in.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationSingle study, not independently replicated
OutcomeA validated creativity measure (AUT, RAT, insight)

Single controlled experiment with a clean crossover design; replication record is thin and mixed.

Wieth, M. B., & Zacks, R. T. (2011). Time of day effects on problem solving: When the non-optimal is optimal. Thinking & Reasoning, 17(4), 387–401. doi:10.1080/13546783.2011.625663 Find this paper ↗

Did you knowThis is the single most counter-intuitive scheduling finding in the literature and it costs nothing to apply. Every organisation puts its innovation workshop in the fresh morning slot. The evidence suggests that is precisely backwards for the generative half.
Watch outDo not schedule the whole day off-peak. Convergence, decision-making and anything analytic should sit in the sharp hours. Split the day if you can.
16Contested

Non-sleep deep rest break

Widely promoted, genuinely pleasant, and resting on a PET study of eight people that never measured creativity.

DRO
20 minAnystate
Full detail

How to run it

A 15–20 minute guided body-scan or Yoga Nidra track between the generation and convergence phases, lying down, eyes closed.

If you run it, be straight with the room about the evidence — which is exactly the teaching moment this activity is best used for.

Facilitator script

  1. 2 min · Frame it honestly
    We are going to do twenty minutes of guided deep rest. I want to be upfront: the famous dopamine finding behind this comes from eight people in a scanner in 2002, and creativity was never measured. We are doing it because a low-demand break is well evidenced, and this is a pleasant way to take one.
  2. 18 min · Run the track
    Lie down, close your eyes, and follow the audio. If you fall asleep that is fine — although the research on sleep-onset suggests you want to stay just this side of it.

The evidence

Kjaer et al. (2002), Cognitive Brain Research. Eight experienced yoga teachers underwent 11C-raclopride PET scans during Yoga Nidra. Binding in the ventral striatum fell 7.9%, which the authors calculate corresponds to a 65% increase in endogenous dopamine release — correlated with a concomitant rise in EEG theta activity. It was the first in vivo demonstration linking endogenous neurotransmitter release to conscious experience, and it is careful work.

It is also n = 8, in expert practitioners, measuring dopamine rather than any cognitive outcome. The step from "dopamine rose" to "you will be more creative" is an inference layered on by others, not a result.

Evidence profile

DesignCorrelational or field-observational
ReplicationSingle study, not independently replicated
OutcomeCreativity was never measured

Small expert-sample PET study. Creativity was never measured — the link to creative output is an inference, not a finding.

Kjaer, T. W., Bertelsen, C., Piccini, P., Brooks, D., Alving, J., & Lou, H. C. (2002). Increased dopamine tone during meditation-induced change of consciousness. Cognitive Brain Research, 13(2), 255–259. Find this paper ↗

Did you knowThe strongest justification for running this is not the dopamine result at all. It is Sio & Ormerod's incubation finding that a low-demand break beats both a high-demand task and doing nothing. NSDR is a well-packaged low-demand break — which is a real, if less exciting, reason to use it.
Watch outDo not present the 65% figure without the sample size. Doing so in front of a room that includes one scientist will cost you the rest of the session.
17Moderate

Device-free nature immersion

A four-day device-free immersion produced a large gain on convergent insight. The 115-minute version below is a scaled-down extrapolation, not the studied intervention — the card says so rather than borrowing the four-day effect size.

DRO
115 minOffsitestate
Full detail

How to run it

Where programme design allows: a multi-day, device-free offsite in a natural setting before a major creative push.

Where it does not: protected, unstructured outdoor time with phones surrendered, framed as work rather than a break. The framing matters — people will not use it properly if it is billed as downtime.

Facilitator script

  1. 5 min · Collect the devices
    Phones in the box. Not on silent — in the box. This is the intervention, not the inconvenience around the intervention.
  2. 90 min · Unstructured outdoor time
    No agenda, no deliverable, no pairing. Walk, sit, whatever you like. If you find yourself solving the problem, let it happen — but do not go looking.
  3. 20 min · Cold capture on return
    Before you talk to anyone, write for five minutes. What surfaced?

The evidence

Atchley, Strayer & Atchley (2012), PLOS ONE. 56 Outward Bound participants on four-to-six-day wilderness expeditions with no electronic devices. 24 took the Remote Associates Test the morning before departure (mean 4.14 of 10); 32 took it on the fourth morning of the trip (mean 6.08) — roughly a 50% improvement. The authors invoke Attention Restoration Theory: modern environments are full of sudden events that hijack attention, while natural settings offer "soft fascination" allowing the executive attention system to replenish.

Evidence profile

DesignCorrelational or field-observational
ReplicationSingle study, not independently replicated
OutcomeA validated creativity measure (AUT, RAT, insight)

Small field study without an active control; outcome is the Remote Associates Test.

Atchley, R. A., Strayer, D. L., & Atchley, P. (2012). Creativity in the wild. PLOS ONE, 7(12), e51474. Find this paper ↗

Did you knowThe researchers ran the pilot on themselves during a five-day backpacking trip in Utah's Grand Gulch before recruiting anyone. An earlier attempt to study attention in nature failed completely — they had sent students into the wild with computers, and the screens attracted moths.
Watch outThis is the weakest design in the moderate tier and you should say so when you teach it. Different people took the test before and during, participants were not randomised, and nature exposure is completely confounded with digital abstinence, exercise, sleep change and social novelty. Plausible and much-loved, but a between-groups comparison in a self-selected sample of 56.

Duration, stated honestly. Atchley and colleagues studied roughly four days of device-free immersion. What you can run inside a programme is the 115 minutes scripted above, and no study has tested that dose. Treat the short version as a plausible extrapolation and do not attach the four-day effect size to it. This card previously advertised 240 minutes while its own script ran to 115; the session builder now counts what you can actually run.
18Moderate

Protecting against fragmentation

Pressure per se is not the killer. Fragmentation is. A hard deadline with protected focus can work; a moderate one with six meetings a day cannot.

DRO
30 minLeadershipclimate
Full detail

How to run it

Audit the innovation calendar. Separate work into exploration blocks — unbounded, protected, no deliverable — and delivery blocks. Never schedule ideation into the gaps between deadlines.

Run it as a live exercise with the actual calendar on screen. It is considerably more confronting that way.

Facilitator script

  1. 10 min · Put the calendar up
    This is last month, real. Point to the block where creative work was supposed to happen. Now count the interruptions inside it.
  2. 12 min · Separate the two kinds
    Colour every block: is this exploration or delivery? Anything that is both is actually delivery.
  3. 8 min · Defend one block
    Pick one exploration block a week and make it genuinely uninterruptible for a month. One. Then we look at what changed.

The evidence

Amabile's componential model and her diary-study research on time pressure found creative thinking was generally lower on high-time-pressure days. The notable exception: pressure did not damage creativity when people felt they were on a meaningful mission, were protected from interruption and distraction, and could focus on a single activity.

Evidence profile

DesignCorrelational or field-observational
ReplicationIndependently replicated at least once
OutcomeReal creative products or field innovation outcomes

Diary study of real projects with real creative output, replicated in related field work; observational, so causal direction is not established.

Amabile, T. M., Hadley, C. N., & Kramer, S. J. (2002). Creativity under the gun. Harvard Business Review, 80(8). Find this paper ↗

Did you knowThe nuance is what makes this usable. A hard deadline with protected focus can produce excellent creative work. A moderate deadline with a fragmented calendar cannot. Most organisations try to fix the deadline and leave the fragmentation untouched.
Watch outDo not let this become an argument for removing all deadlines. Unbounded exploration with no endpoint produces its own failure mode, and teams know it.
19Contested

The 30-second eye movement claim

A real study, a real journal, a mechanism quietly falsified downstream. Run it as a source-tracing exercise rather than a technique.

DRO
25 minAnystate
Full detail

How to run it

You will meet this one in the wild: thirty seconds of looking left and right boosts creativity by connecting the two sides of your brain. The underlying paper exists and is properly published. Almost everything added to it in transmission does not hold.

Rather than running the technique, run the trace. Hand out the 2009 abstract and the 2015 replication abstract. Ask the room to identify, in fifteen minutes, exactly which claim was made by whom — and where the sentence "connects the two hemispheres" actually entered the story.

Facilitator script

  1. 3 min · State the popular version
    Here is a claim you will have seen on LinkedIn. Thirty seconds of horizontal eye movements increases your ideation, because it connects the two sides of your brain. Hands up if that sounds plausible.
  2. 7 min · Read the original abstract
    This is the actual 2009 paper. Read it and tell me three things: who benefited, on which measures, and whether the hemisphere explanation is presented as a finding or as a hypothesis.
  3. 8 min · Read the replication
    Now the 2015 paper. Note the design — proponents and sceptics agreed the method together, preregistered it, and committed to publish either way. What did they find, and what did it do to the mechanism?
  4. 7 min · Generalise the lesson
    The paper was fine. The journal was fine. The failure happened entirely in transmission, between the abstract and the slide. That is the pattern you should now expect by default — and the reason every card on this site names its primary source.

The evidence

Shobe, Ross & Fleck (2009), Brain and Cognition, tested 62 participants on the Alternate Uses Task after a 30-second bilateral eye movement task versus a central-fixation control. Bilateral eye movements raised originality and categorical distinctiveness in strong-handers only; mixed-handers showed no effect from the manipulation. Fluency, appropriateness and detail were unaffected. So the honest summary is that two of five sub-scores improved, in roughly half the sample.

The interhemispheric interaction account was always the proposed mechanism, not a result. It has since taken heavy damage. Matzke et al. (2015) ran a preregistered adversarial collaboration — proponents and sceptics agreed the design in advance and committed to publish regardless — and failed to replicate the eye-movement benefit on memory, with strong Bayesian evidence for the null. Roberts, Fernandes & MacLeod (2020) found the same pattern in PLOS ONE: a weak replication in Experiment 1, then substantial Bayesian evidence for a null in a larger, better-powered Experiment 2. Samara et al. (2011) had already found no consistent change in interhemispheric EEG coherence following bilateral saccades across six frequency bands.

None of that formally overturns the creativity result, which was a different task. It does dismantle the explanation everyone repeats.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationFailed a preregistered replication
OutcomeA validated creativity measure (AUT, RAT, insight)

Competent original experiment whose proposed mechanism failed a preregistered adversarial replication.

Shobe, E. R., Ross, N. M., & Fleck, J. I. (2009). Influence of handedness and bilateral eye movements on creativity. Brain and Cognition, 71(3), 204–214. See also Matzke et al. (2015), JEP: General, 144(1), e1–e15. Find this paper ↗

Did you knowThe 2015 replication is worth teaching for its method alone. Two research groups who disagreed — one believing the effect was real, one believing it was not — jointly designed the study, published an adversarial collaboration agreement in advance, and committed to submitting whatever came out. It is one of the cleanest examples in psychology of how to settle a dispute rather than prolong it.
Watch outIf you want an eye-based intervention that has held up better, close them rather than move them. Ritter, Abbing & van Schie (2018) found eye-closure improved both divergent and convergent creativity with no effect on working memory — suggesting specificity to creativity rather than a general cognitive boost. See activity 14.
20Moderate

The decontextualised brief

Strip every industry noun out of the brief before the machine sees it. Context is what stops you finding the answer somewhere else.

DRO
30 min2–12frame
Full detail

How to run it

Write the brief twice. Version A is how you would normally phrase it, full of your domain vocabulary. Version B has every industry-specific noun removed and every assumption about the solution stripped out — a pure functional statement.

Run both through the same model. Compare the return sets side by side, and count how many results in each set come from outside your own sector.

Facilitator script

  1. 5 min · Write version A
    Write the brief the way you would send it to a colleague. Jargon welcome. This is your control condition.
  2. 8 min · Strip it
    Now remove every noun that names your industry, your product, or your existing solution. If someone outside your sector could not tell what business you are in, you have done it right. Keep the function, lose the furniture.
  3. 10 min · Run both
    Same model, same settings, both briefs. Do not clean up the results yet.
  4. 7 min · Count the outsiders
    How many results in each set come from outside your sector? That difference is the cost your own vocabulary has been quietly imposing on every search you have ever run.

The evidence

Hassan & Nair (2026), Research Policy, identify de-contextualizing as one of three AI complementarities in innovation search, derived inductively from 41 interviews across 27 organisations. Participants described AI as returning options they had already ruled out on assumed grounds — one R&D manager reported selecting two technologies that human search had dismissed on feasibility assumptions. Another noted that AI has no gut feeling, so it never pre-dismisses; a human "just neglects" an option and then "might miss something important."

The authors frame the human–AI pairing explicitly as divergent and convergent thinking, quoting one participant: AI helps you go broader more quickly, but at some point you have to converge — and AI is much less useful on that half.

Evidence profile

DesignQualitative, case-based or theoretical only
ReplicationSingle study, not independently replicated
OutcomeCreativity was never measured

Framework-grade. Inductive interview study; no outcome was measured and no comparison was run.

Hassan, R., & Nair, S. (2026). Searching together: Human–AI complementarities in innovation search. Research Policy, 55, 105597. Open access, CC BY. Find this paper ↗

Did you knowThe same paper reports AI surfacing old Russian and East German patents through machine translation that a specialist said he would "never ever" have found manually — and a Bluetooth beacon system from airport passenger tracking, retrieved for a problem in a completely unrelated sector.
Watch outThis is framework-grade evidence, not effect-grade. The study is inductive and qualitative, no outcome was measured, and 25 of the 27 organisations were recruited through the AI vendor's own customer network. Teach the mechanism; do not quote a number.
21Moderate

AI as the outsider seat

Activity 09 with a machine in the chair. Same mechanism as the InnoCentive finding — distance from the domain is the asset.

DRO
40 min4–15diverge
Full detail

How to run it

Run the outsider seat protocol, but assign the role to the search tool. Give it the decontextualised brief from activity 20, and require it to return results from at least five named sectors none of which is yours.

Then — and this is the part teams skip — assign a human translator whose only job is to map each foreign result back to a mechanism in your domain. Without the translator you get a list of curiosities.

Facilitator script

  1. 5 min · Name the forbidden sector
    Your own industry is off the table for this round. Anything the tool returns from inside our sector, we discard, however good it looks.
  2. 15 min · Retrieve wide
    Five sectors minimum, none of them ours. Do not filter for relevance yet — filtering is the next phase and doing it now collapses the whole exercise.
  3. 15 min · Translate, do not admire
    For each result, one sentence: what is the mechanism, and what is its equivalent here? Not "that is interesting" — a mechanism and a mapping.
  4. 5 min · Debrief the discomfort
    Which results did you want to throw away fastest? Note them. That instinct is the thing this activity exists to interrupt.

The evidence

Hassan & Nair (2026) identify boundary-spanning search as AI's primary complementary capability — crossing regional, linguistic, disciplinary and organisational boundaries. Participants described AI search as "not based on experience, it is based on what is available," and noted it "doesn't care where it comes from" geographically.

The mechanism is the same one Jeppesen & Lakhani (2010) measured quantitatively across 166 InnoCentive challenges and 12,000+ solvers: the probability of a winning solution rose with the distance between the solver's expertise and the problem domain. That study measured outcomes; this one describes a machine doing the same job. Use the older paper for the evidence and the newer one for the practice.

Evidence profile

DesignCorrelational or field-observational
ReplicationSingle study, not independently replicated
OutcomeAn adjacent cognitive outcome only

Framework-grade for the AI half; the underlying distance-and-solving effect (activity 09) is where the outcome evidence lives.

Hassan, R., & Nair, S. (2026). Research Policy, 55, 105597. Read alongside Jeppesen, L. B., & Lakhani, K. R. (2010), Organization Science, 21(5), 1016–1033. Find this paper ↗

Did you knowThe same paper documents AI finding knowledge the organisation already owned. One vendor described telling a client about a patent filed three years earlier — by the client's own colleague. "Well, you should talk to him." Large multinationals, the participant noted, routinely invent something they later need and forget they have it.
Watch outBoundary-spanning without translation produces a longer list, not a better one. One participant in the study describes assessment and translation of AI output as "really time consuming" — the authors themselves question whether the cost savings survive it.
22Contested

The hallucination harvest

Treat the irrelevant returns as need–solution pairs looking for a problem. The strongest examples are real; the general claim is not yet tested.

DRO
35 min3–12diverge
Full detail

How to run it

Do not delete the discard pile. Take the results your search tool returned that were clearly wrong, irrelevant, or fabricated, and run them through one question: if this were the answer, what would the question have been?

Most produce nothing. The exercise costs twenty minutes and occasionally produces a problem nobody had articulated.

Facilitator script

  1. 5 min · Recover the discards
    Everything we threw out in the last hour — put it back on the table. Yes, including the ones that were obviously nonsense.
  2. 15 min · Reverse the pair
    For each one: if this were a good answer, what problem would it be answering? Write the problem, not a critique of the answer.
  3. 10 min · Test against reality
    Do any of those invented problems actually exist for us? Even faintly? Circle them.
  4. 5 min · Set expectations honestly
    Most sessions produce nothing here and that is the expected result. This is a lottery ticket with a twenty-minute price, not a method.

The evidence

Hassan & Nair (2026) argue that AI hallucination, used deliberately, can seed serendipitous discovery — connecting it to Hippel & Von Krogh's (2016) concept of need–solution pairs, where a viable solution reveals a problem nobody had formulated. The strongest supporting cases in the paper are real and published: de novo protein design by deep network hallucination (Anishchenko et al., 2021, Nature), and the Robot Scientist Eve, which while screening triclosan against one malarial enzyme target unexpectedly identified a different one, DHFR (Bilsland et al., 2018, Scientific Reports).

Those are genuine discoveries. What does not yet exist is any evidence that deliberately harvesting hallucinations works as a repeatable workshop method for ordinary problems.

Evidence profile

DesignQualitative, case-based or theoretical only
ReplicationFailed a preregistered replication
OutcomeCreativity was never measured

The weakest card in the library. Individual examples are documented; the general claim has never been tested.

Hippel, E. von, & von Krogh, G. (2016). Identifying viable need–solution pairs. See also Anishchenko, I., et al. (2021), Nature, 600(7889), 547–552. Find this paper ↗

Did you knowAn economist quoted in the paper puts the case in one line: a hallucination is "a kind of thing that looks true" — which is exactly what a hypothesis is. The difference between a hallucination and a hypothesis is whether anyone bothers to test it.
Watch outGraded contested deliberately. The published cases are in protein design and automated drug screening, where a hypothesis can be tested cheaply and at scale by machine. Your workshop cannot test a hundred thousand candidates overnight. Do not let this become a licence to treat confident nonsense as insight.
23Moderate

Guarding the AI shortlist

The converge phase is where AI-sourced novelty dies — and the people killing it will sincerely report that they wanted it.

DRO
40 minDecision-makersconverge
Full detail

How to run it

Two rules, both structural.

One: review the full return set, never the pre-filtered shortlist. Whoever filtered it applied their own contextual assumptions, which is exactly what the AI search existed to bypass.

Two: score novelty blind to source. Strip every result of whether it came from AI or human search before anyone rates it. Reveal the source only after scoring.

Facilitator script

  1. 5 min · Demand the long list
    I do not want the shortlist. I want everything the search returned, including what was cut and who cut it.
  2. 12 min · Strip and shuffle
    Remove all provenance. Nobody rating these should know which came from the tool and which from our own network.
  3. 13 min · Score novelty only
    Novelty first, on its own, by people who are not scoring feasibility. If you catch yourself thinking "we could never do that", you are on the wrong axis.
  4. 10 min · Force the advance and reveal
    One high-novelty concept advances with a named advocate. Now reveal the sources — and notice whether the AI-sourced items clustered in what you rejected.

The evidence

Hassan & Nair (2026) position human judging as the corrective that filters AI output and mitigates AI bias. That is the paper's weakest link, because it is never tested against Mueller, Melwani & Goncalo (2012), who showed that under uncertainty people develop an implicit bias against creativity that impairs their ability to recognise a creative idea — while sincerely reporting that they want one. Novel outputs from unfamiliar domains with uncertain relevance is a precise description of the uncertainty condition that triggers it.

Second reason for structure: Doshi & Hauser (2024), Science Advances. In an online experiment, writers given GPT-4 story ideas produced work rated more creative, better written and more enjoyable — with the largest gains among less creative writers, effectively equalising evaluations. But the AI-assisted stories were measurably more similar to each other than the human-only stories. The authors describe it as a social dilemma: individually better off, collectively a narrower range of novel content.

Evidence profile

DesignQualitative, case-based or theoretical only
ReplicationSingle study, not independently replicated
OutcomeCreativity was never measured

Framework-grade. Derived from interviews, with no measured effect on selection quality.

Mueller, J. S., Melwani, S., & Goncalo, J. A. (2012). Psychological Science, 23(1), 13–17. Doshi, A. R., & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances, 10(28), eadn5290. Find this paper ↗

Did you knowDoshi & Hauser cuts against the whole boundary-spanning argument, and it appears in the Hassan & Nair paper as a single sentence in the final paragraph. If every organisation runs the same decontextualised search against the same models, the distance advantage collapses into convergence. The tool that widens your search may narrow everyone's.
Watch outIf your team reviews a shortlist someone else filtered, you have not run this activity — you have ratified another person's contextual assumptions with extra steps.
24Strong

The fixation demonstration

Show half the room an example before they ideate. They will reproduce its flaws — including the flaws you warned them about. The best-replicated finding in design cognition.

DRO
35 min6–40frame
Full detail

How to run it

Split the room in two, out of earshot. Group A gets the design brief alone. Group B gets the brief plus one worked example, and — this is the part that makes the demonstration land — an explicit instruction that the example contains three named flaws which they must avoid.

Both groups ideate for the same time. Then post the outputs together and count how many of Group B's ideas carry each named flaw. Historically, most of them do.

Facilitator script

  1. 3 min · Split and brief separately
    I am going to give the two halves of the room slightly different instructions. Do not compare notes until we post the results — the whole exercise depends on it.
  2. 12 min · Ideate
    Twelve minutes, sketch or write as many distinct designs as you can. Group B, remember: the example has three flaws, listed at the bottom. Avoid them.
  3. 10 min · Post and count
    Everything on the wall, Group A on the left, Group B on the right. Now we count. How many designs on the right contain flaw one? Flaw two? Flaw three?
  4. 10 min · Debrief the mechanism
    Group B were warned. In writing. And it made almost no difference, because fixation is not a failure of attention — it is what happens when an example becomes the category you search inside. Now: what is the example that is already sitting in your team's head before every kickoff?

The evidence

Jansson & Smith (1991), Design Studies, gave engineering designers a brief with or without an example solution. Those shown the example reproduced its features — including features they had been explicitly told were defective — at markedly higher rates. The authors named the effect design fixation.

It is among the most robustly replicated results in design cognition. Purcell & Gero (1996) replicated it and showed it varies by discipline; Chrysikou & Weisberg (2005) replicated it in a psychology sample and showed that a defixating instruction can partly reduce it; Linsey et al. (2010) demonstrated it in expert faculty engineers, not just students, and tested countermeasures.

Two things make this stronger evidence than most of the library. The outcome is real design output rather than an Alternate Uses score, and the effect appears in experts, which is where most creativity findings quietly disappear.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationMeta-analytic or consistently replicated
OutcomeReal creative products or field innovation outcomes

Controlled between-subjects experiment, replicated across four decades and multiple domains, scored on real design output.

Jansson, D. G., & Smith, S. M. (1991). Design fixation. Design Studies, 12(1), 3–11. doi:10.1016/0142-694X(91)90003-F Find this paper ↗

Did you knowWarning people about the flaws does not reliably work, but the timing of the example does. Present examples after a first independent generation round and the fixation cost largely disappears — which is why the running order of a workshop is not an administrative detail.
Watch outDo not conclude that examples are bad. Agogué et al. (2014) showed examples can also raise originality — it depends entirely on whether the example sits inside or outside the obvious solution path. See activity 25, which is the other half of this finding.
25Strong

Expansive versus restrictive examples

The other half of fixation. An example inside the obvious solution path suppresses originality; an example outside it raises originality. Same intervention, opposite sign.

DRO
40 min8–40frame
Full detail

How to run it

This is the operational core of C–K design theory (see activity 26) reduced to something you can run in forty minutes.

Take one brief. Prepare two example sheets. The restrictive sheet shows a solution from inside the obvious path everyone already searches. The expansive sheet shows a solution from a path almost nobody considers. Assign the room at random — genuinely at random, count off — and give each half one sheet. Ideate, then score originality against the mapped solution space rather than by gut feeling.

The classic vehicle is the egg task: ensure a hen's egg dropped from ten metres does not break. Almost everyone searches the same region — wrap it, cushion the landing, slow the fall. The expansive paths are the ones that change the egg, change the ground, or change what "not break" means.

Facilitator script

  1. 6 min · Map the solution space first
    Before we ideate, we map. On the board: what are the families of solution everyone reaches for? Cushion it, slow it, protect it. Those are the paths inside the obvious set. Anything not on this board is the expansive region.
  2. 4 min · Randomise and distribute
    Count off, ones and twos. Ones take sheet A, twos take sheet B. Do not read your neighbour's sheet.
  3. 12 min · Generate
    Twelve minutes. As many distinct solutions as you can. One per card.
  4. 10 min · Score against the map
    Sort every card onto the map we drew. Anything landing outside the mapped families counts as expansive. Count the two groups separately.
  5. 8 min · Reveal and generalise
    You had different sheets. Group A's example sat inside the obvious set; Group B's sat outside it. Look at the counts. The lesson is not that examples help or hurt — it is that the example silently tells people where the search space ends.

The evidence

Agogué, Kazakçi, Hatchuel, Le Masson, Weil, Poirel & Cassotti (2014), Journal of Creative Behavior, ran the egg task with restrictive and expansive example conditions. Restrictive examples — those lying within the commonly explored solution path — depressed originality relative to no example. Expansive examples — lying outside it — raised originality. The manipulation was the position of the example in the solution space, not its quality or detail.

The finding has an unusually good follow-up record for this literature. Agogué, Poirel, Pineau, Houdé & Cassotti (2014), Thinking Skills and Creativity, replicated the pattern across age groups and tested training. Agogué, Le Masson, Dalmasso, Houdé & Cassotti (2015), Psychology of Aesthetics, Creativity, and the Arts, found professional industrial designers and engineers resist classical solutions differently from novices. Ezzat et al. (2020) separated the specificity of an example from its abstraction and found opposing effects.

What raises this above a single cute result is that originality is scored against an independently mapped solution space rather than by rater impression, which removes a large chunk of the subjectivity that weakens most divergent-thinking scoring.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationIndependently replicated at least once
OutcomeA validated creativity measure (AUT, RAT, insight)

Controlled experiment with random assignment, independently replicated in adults, children and professional designers; originality scored against a mapped solution space.

Agogué, M., Kazakçi, A., Hatchuel, A., Le Masson, P., Weil, B., Poirel, N., & Cassotti, M. (2014). The impact of type of examples on originality. The Journal of Creative Behavior, 48(1), 1–12. doi:10.1002/jocb.37 Find this paper ↗

Did you knowMapping the solution space before ideation is doing half the work of C–K theory without the notation. Once a room can see where the obvious paths stop, "be more creative" becomes a specific instruction: generate in the unshaded region.
Watch outThe mapping step is the one facilitators skip, and skipping it collapses the activity into ordinary brainstorming. If you have not mapped the obvious paths in advance, you cannot tell an expansive idea from an enthusiastic one.
26Moderate

C–K mapping: concept and knowledge

A formal design theory that gives the frame phase a notation. Framework-grade evidence, not effect-grade — but its fixation predictions have been tested experimentally, which is rare.

DRO
60 min4–20frame
Full detail

How to run it

C–K theory, developed by Hatchuel and Weil at MINES Paris, models designing as the joint expansion of two spaces. K-space holds propositions with a known logical status — things you can call true or false. C-space holds undecidable propositions: "a boat that has no hull" is neither true nor false given current knowledge, which is precisely what makes it a concept rather than a fact.

Designing is the four operators between them: C→C (partitioning a concept into more specific ones), C→K (asking what knowledge a concept demands), K→C (using knowledge to open a new concept), K→K (ordinary learning). A design converges when a concept acquires a logical status and crosses into K.

Run it on the wall. Two colours: one for K, one for C. Start with the brief as a C-proposition. Every time someone says "but we know that…", that is a K card. Every time a K card makes a new undecidable proposition available, draw the arrow.

Facilitator script

  1. 8 min · Write the initial concept
    State the brief as a proposition we cannot yet call true or false. Not 'improve the onboarding' — that is a goal. Something like 'an onboarding with no first session'. If someone can immediately tell me it is true or false, it is not a concept, it is knowledge.
  2. 12 min · Harvest K
    What do we actually know that bears on this? Constraints, results, prior attempts, regulations. Blue cards. Include the things we know that we wish were not true.
  3. 15 min · Partition the concept
    Now split the concept into narrower ones. Two kinds of split: restrictive, which adds a property from inside what we already do, and expansive, which adds a property that breaks an assumption in K. Mark every expansive partition with a star. Most teams produce almost none, and that is the finding.
  4. 15 min · Follow the starred branches
    For each starred branch: what would we have to know for this to become decidable? That question is the C→K operator, and the answer is your actual research agenda.
  5. 10 min · Close the loop
    Which new knowledge, once acquired, would open a concept we could not previously state? That is where a design programme comes from, as opposed to a list of ideas.

The evidence

Hatchuel & Weil (2009), Research in Engineering Design, give the formal statement of C–K theory: design as the co-expansion of a Concept space of undecidable propositions and a Knowledge space of decidable ones, connected by four operators. The theory's distinctive claim is that creativity is not a psychological trait but a logical operation — expansive partitioning — which can be taught and inspected.

The evidence has two very different halves, and the honest thing is to keep them apart.

The theory itself is framework-grade. Its support is industrial case studies and action research — KCP workshops at Renault, Thales, RATP and others, documented by Le Masson, Weil & Hatchuel. No randomised comparison against an active control has tested whether C–K sessions produce more or better innovation than another structured method. Nobody has measured the outcome. Teach the mechanism; do not quote a number.

Its fixation predictions are effect-grade. The expansive-versus-restrictive prediction has been tested experimentally and replicated — see activity 25. That is a genuine and unusual asset: most design methodologies in circulation have no experimental arm at all.

Evidence profile

DesignQualitative, case-based or theoretical only
ReplicationSingle study, not independently replicated
OutcomeCreativity was never measured

Framework-grade. A formal design theory supported by industrial case studies and action research; the theory itself has not been tested against an active control for creative output.

Hatchuel, A., & Weil, B. (2009). C–K design theory: an advanced formulation. Research in Engineering Design, 19(4), 181–192. See also Le Masson, Weil & Hatchuel, Design Theory (Springer, 2017). doi:10.1007/s00163-008-0043-4 Find this paper ↗

Did you knowThe undecidability test is the most useful thing to take away even if you never draw a C–K diagram. If a room can immediately tell you a proposition is true or false, it is not a concept and it will not expand anything. Most kickoff meetings run entirely in K-space and then wonder why the output is incremental.
Watch outTwo failure modes. First, the notation seduces: teams produce a beautiful tree and no new knowledge. The C→K arrow is the load-bearing one, not the tree. Second, do not present C–K to a room as evidence-backed in the way activity 24 is. It is a formalism with case support and one tested prediction — say so, or you are committing exactly the transmission error section 05 describes.
27Moderate

Structured dissent

The finding that contradicts the first rule of brainstorming. Instructed debate and criticism produced more ideas than the classic defer-judgment rule — and this site owes you the tension.

DRO
45 min5–25diverge
Full detail

How to run it

Run the same problem under two conditions with two groups. Condition A gets the standard rule: no criticism, defer judgment. Condition B gets an explicit instruction to debate and to challenge each other's ideas as they arrive.

Count total ideas and, separately, ideas generated afterwards when both groups are asked individually for anything else that occurs to them. The second count is where the interesting difference usually shows up.

Facilitator script

  1. 5 min · Set the two rules
    Group A: the classic rule. Nothing is criticised, all ideas are welcome, judgment is deferred. Group B: debate as you go. If you think an idea is wrong, say so and say why. This is not permission to be unpleasant — it is an instruction to engage.
  2. 20 min · Run both
    Twenty minutes on the same problem. Capture everything, including things that were argued down.
  3. 10 min · The individual tail
    Now, alone and in silence: anything else that has occurred to you since. Five minutes, write it down. This is a separate count.
  4. 10 min · Compare and complicate
    Two numbers per group: in-session ideas and tail ideas. Then the harder question — which room would you rather have been in, and is that the same as which room produced more?

The evidence

Nemeth, Personnaz, Personnaz & Goncalo (2004), European Journal of Social Psychology, compared brainstorming instructions with debate instructions in samples from the United States and France. The debate condition produced more ideas, and the advantage persisted into the individual ideas participants reported afterwards. The finding sits in a longer programme on minority dissent, in which exposure to a persistently argued minority view broadened rather than narrowed search — measured on word-association and problem-solving tasks, not on self-report.

This is a direct challenge to Osborn's defer-judgment rule, which activity 01 uses. Both cannot be straightforwardly right, and it would be dishonest for this library to carry one and quietly omit the other.

The reconciliation most defensible on current evidence: the harm in ordinary group brainstorming is production blocking and evaluation apprehension, not disagreement as such (see activity 07). Structured, task-focused dissent with everyone still able to produce appears to help; unstructured social evaluation does not. That reading is plausible and partly supported — it is not settled.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationIndependently replicated at least once
OutcomeA validated creativity measure (AUT, RAT, insight)

Controlled laboratory experiments with random assignment, replicated across cultures; outcomes are validated ideation and association measures rather than field innovation.

Nemeth, C. J., Personnaz, B., Personnaz, M., & Goncalo, J. A. (2004). The liberating role of conflict in group creativity. European Journal of Social Psychology, 34(4), 365–374. doi:10.1002/ejsp.210 Find this paper ↗

Did you knowNemeth's earlier work found that exposure to a wrong but consistently argued minority position improved performance on unrelated tasks. The value of dissent in those studies did not depend on the dissenter being right.
Watch outDissent and psychological safety are not opposites, but a room without the second cannot survive the first. Do not run this activity with a group that has not done activity 10, and never with a group where a status difference means the debate will only run downward.
28Moderate

Sleeping on it

Overnight consolidation, not the fifteen-second version. Structure the session across two days and pose the problem before people leave.

DRO
25 minAnystate
Full detail

How to run it

The mechanism is not "take a break." It is that sleep reorganises recently encoded material, and insight sometimes falls out of that reorganisation. Which means the scheduling instruction is specific: the problem has to be encoded before sleep, in enough depth that there is something to consolidate.

Day one: pose the problem, work it hard enough that everyone has hit a wall, then stop — explicitly, with no request to think about it overnight. Day two: first thing, individually and in silence, before any discussion.

Facilitator script

  1. 10 min · Encode properly, day one
    Ten minutes on the problem, and I want you to get stuck. Being stuck is the condition, not the failure. Write down where the wall is.
  2. 2 min · Close it down
    We stop here. I am deliberately not asking you to think about this tonight — the effect does not require deliberate rumination and asking for it tends to produce anxiety instead.
  3. 8 min · Day two, silent first
    Before anyone speaks: eight minutes, alone, on the same problem. If something arrived overnight, it will not survive the first three minutes of group conversation, so write it before we talk.
  4. 5 min · Collect the overnight set
    Hands up if something shifted. Now compare where the wall was yesterday to where it is now — the interesting cases are where the wall moved rather than where an answer appeared.

The evidence

Wagner, Gais, Haider, Verleger & Born (2004), Nature, trained participants on a number-reduction task containing a hidden shortcut rule. Participants who slept between training and retest discovered the rule at more than twice the rate of matched groups who stayed awake for equivalent periods, day or night. The comparison against a wake-control of the same duration is what makes it more than a rest effect.

The wider sleep-and-memory consolidation literature is strong and heavily replicated. The specific leap from consolidation to creative insight is the softer link — the number-reduction task is a rule-discovery task, and generalising it to open-ended ideation is an inference the study does not make.

Read alongside activity 06, which uses the very different sleep-onset (N1) window, and activity 04, whose incubation meta-analysis covers the waking case.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationIndependently replicated at least once
OutcomeA validated creativity measure (AUT, RAT, insight)

Controlled experiment with matched wake-control groups; replicated in related paradigms, though the specific insight result has been contested on task-specificity grounds.

Wagner, U., Gais, S., Haider, H., Verleger, R., & Born, J. (2004). Sleep inspires insight. Nature, 427, 352–355. doi:10.1038/nature02223 Find this paper ↗

Did you knowThe instruction not to think about it overnight is doing real work. Deliberate rumination during a break suppresses the incubation benefit in Sio & Ormerod's meta-analysis — the low-demand condition is the one that helps.
Watch outThis is the single easiest finding in the library to over-sell, because everybody already believes it. Sleeping on a problem you have not properly encoded does nothing at all, and a two-day workshop is not automatically better than a one-day workshop.
29Moderate

Open-monitoring meditation

Not meditation in general — one specific style. Open-monitoring raised divergent thinking; focused-attention did not, and in some samples reduced it.

DRO
30 minAnystate
Full detail

How to run it

The distinction is the whole activity. Focused-attention meditation trains sustained attention on a single object — breath, mantra, a point. Open-monitoring trains non-reactive awareness of whatever arises, without selecting.

Run both, ideally with the room split, and put an Alternate Uses Task afterwards. Then discuss why a room that assumes "mindfulness makes you creative" would have predicted no difference — and what it costs to be that imprecise.

Facilitator script

  1. 4 min · Explain the split, not the prediction
    Two styles. One narrows attention onto a single object. One widens it and refuses to select. Do not tell each other which you think will help — I will ask afterwards.
  2. 12 min · Run the practice
    Twelve minutes. Group A, attention on the breath; when it wanders, return it. Group B, notice whatever arises — sound, sensation, thought — and do not follow any of it, and do not push any of it away.
  3. 8 min · Immediate AUT
    Straight into it, no discussion. Eight minutes: as many uses for a paperclip as you can. Quantity.
  4. 6 min · Score and discuss
    Fluency counts first, then originality. Then the real question: if the two styles differ, what does that do to every claim you have heard that begins 'mindfulness improves…'?

The evidence

Colzato, Ozturk & Hommel (2012), Frontiers in Psychology, compared open-monitoring and focused-attention meditation within experienced practitioners and found open-monitoring improved divergent thinking on the Alternate Uses Task, while focused-attention did not — and if anything favoured convergent performance. Lippelt, Hommel & Colzato's (2014) review of the two styles reports a broadly consistent dissociation across studies.

The caution is real. Sample sizes in this literature are small, several studies use experienced meditators rather than naïve participants — which limits what a single workshop session can claim — and the broader meditation-and-cognition field has a well-documented publication-bias problem. Treat the dissociation as the finding worth teaching, not the magnitude.

Set against activity 16: this is better evidenced than the NSDR protocol on that card, and it is the version to run if you want a state intervention with an actual creativity outcome attached.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationIndependently replicated at least once
OutcomeA validated creativity measure (AUT, RAT, insight)

Randomised laboratory comparison of two meditation styles against each other, replicated in independent samples; outcome is the Alternate Uses Task.

Colzato, L. S., Ozturk, A., & Hommel, B. (2012). Meditate to create: the impact of focused-attention and open-monitoring training on convergent and divergent thinking. Frontiers in Psychology, 3, 116. doi:10.3389/fpsyg.2012.00116 Find this paper ↗

Did you knowThe dissociation is a good diagnostic for any wellbeing claim brought into a creativity programme. Ask which sub-process the intervention is supposed to move. "It reduces stress" is not an answer, because activity 18 shows that stress and creative output have a more complicated relationship than the claim assumes.
Watch outTwelve minutes of open-monitoring practice in a room of complete beginners is not the same intervention that was studied, and some participants find it genuinely uncomfortable. Make it opt-out without comment, and do not present a single session as a demonstrated effect.
30Moderate

Design Heuristics cards

An empirically derived prompt deck, not an invented acronym. The heuristics were extracted from what expert designers actually did — which is why they outperform SCAMPER.

DRO
45 min2–30diverge
Full detail

How to run it

Give each participant three heuristic cards drawn at random — add motion, change the surface, make it hollow, use a common base, and so on. Each card is a single transformation with a worked visual example.

One sketch per card, then swap decks and repeat. The rule that makes it work: the heuristic is not optional. If the card says change the surface, the sketch must change the surface, however unpromising that seems.

Facilitator script

  1. 5 min · Frame the deck honestly
    These are not somebody's clever acronym. They were derived by analysing what expert designers and award-winning products actually did, then reduced to the transformations that recurred. That provenance is why we are using this deck rather than a more famous one.
  2. 15 min · First pass, three cards
    Three cards each, one sketch per card, five minutes per sketch. You must apply the card even if you think it is wrong for this problem. Bad forced sketches are the point — they move you out of your first solution.
  3. 15 min · Swap and repeat
    Pass your cards two seats left. Three more sketches. Notice whether the second set is easier or harder than the first.
  4. 10 min · Sort by distance
    Lay everything out. Which sketches are variations on where you started, and which are somewhere genuinely else? Count the second group per person. That count is what the deck is for.

The evidence

Design Heuristics were derived empirically by Yilmaz, Seifert, Daly and Gonzalez from protocol studies of practising designers, analysis of award-winning products and expert interviews — rather than being invented and then promoted. The resulting set of 77 heuristics has been tested in controlled studies with engineering and industrial design students, with concepts rated by independent judges for novelty and quality; heuristic-supported groups produced more varied and more novel concepts than unsupported controls.

Two reasons this earns a place over the better-known idea-prompt methods. The outcome measured is rated design output rather than an Alternate Uses score. And the tool's content has a documented derivation, so a facilitator can answer "where did these prompts come from" with something other than "a book."

The limits are ordinary but worth stating: student samples predominate, judge-rated novelty carries the usual subjectivity, and the comparison is often against no support rather than against another prompt method — so the deck beating nothing is better established than the deck beating SCAMPER.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationIndependently replicated at least once
OutcomeReal creative products or field innovation outcomes

Controlled studies with random assignment in engineering and industrial design samples; outcomes are rated design concepts, not lab proxies.

Yilmaz, S., Daly, S. R., Seifert, C. M., & Gonzalez, R. (2016). Evidence-based design heuristics for idea generation. Design Studies, 46, 95–124. doi:10.1016/j.destud.2016.05.001 Find this paper ↗

Did you knowThe forcing rule matters more than the deck. A prompt that participants are allowed to reject becomes a prompt they reject whenever it would have moved them, which is exactly when it was working.
Watch outCompare with activity 34. The reason this card is graded higher than SCAMPER is not that its prompts are cleverer — several overlap almost exactly — but that somebody did the studies. If a comparable trial someday favours SCAMPER, this card should move.
31Moderate

Ask the creators, not the managers

Who forecasts a novel idea's success best is not who you think, and it is not the person whose job title says decision-maker.

DRO
40 min6–30converge
Full detail

How to run it

Take a genuine shortlist — ideas you have not yet decided between. Have three groups forecast success independently and in private: the people who created the ideas, the people who manage the portfolio, and if you can get them, people from the intended audience.

Do not average them into a single number. Compare the shape of the three rankings, and specifically look at which ideas the manager group ranks lowest that another group ranks high. Those are the ideas your normal process kills.

Facilitator script

  1. 5 min · Assign roles honestly
    Three sheets. You are scoring as creator, as manager, or as audience. If you happen to be both a creator and a manager here, score as creator and hand your manager hat to someone else — mixing the two is what we are trying to see.
  2. 12 min · Private forecasting
    Score every idea for likely success, alone, no discussion. Private is not a formality — one confident voice collapses the whole comparison.
  3. 13 min · Three rankings, side by side
    Post all three. Do not average. I want to see the disagreements, because the disagreements are the finding.
  4. 10 min · Find the killed ideas
    Which ideas do managers rank in the bottom third and creators or audience rank in the top third? Name them. In your normal process, every one of those is dead by Thursday.

The evidence

Berg (2016), Administrative Science Quarterly, had creators, managers and audience members forecast the success of novel ideas, then compared those forecasts against real downstream outcomes. Creators were better forecasters of novel ideas than managers on the measures used — the reverse of the assumption embedded in most stage-gate processes. The proposed mechanism is a divergent-then-convergent asymmetry: the same mindset that generates a novel idea also supports evaluating novelty, while a purely evaluative mindset systematically discounts it.

Read with activity 11. Mueller, Melwani & Goncalo's bias against creativity explains why the evaluative role penalises novelty; Berg's work says something about who to ask instead.

The evidence limit is clear. The role comparison is observational — people were not randomly assigned to be creators or managers — so selection is a live alternative explanation, and the finding has not been independently replicated. What earns its place is that the outcome is real success data rather than a lab proxy.

Evidence profile

DesignCorrelational or field-observational
ReplicationSingle study, not independently replicated
OutcomeReal creative products or field innovation outcomes

Field study with real innovation outcomes and pre-registered-style forecasting design, but observational role comparison rather than random assignment; not yet independently replicated.

Berg, J. M. (2016). Balancing on the creative highwire: forecasting the success of novel ideas in organizations. Administrative Science Quarterly, 61(3), 433–468. doi:10.1177/0001839216642211 Find this paper ↗

Did you knowThe practical version costs nothing. Before a selection meeting, collect private forecasts from the people who generated the ideas, and put those rankings in the room next to the decision-makers'. You are not obliged to follow them — you are obliged to notice the gap.
Watch outThis is not an argument that creators should decide. Creators over-rate their own work in other studies, and the forecasting advantage in this one is about novel ideas specifically. Use it to add a ranking, not to remove a decision-maker.
32Moderate

Improvisation training

The corporate staple that turns out to have a small controlled literature behind it — with the benefit landing on divergent thinking and affect, not on collaboration.

DRO
50 min6–24climate
Full detail

How to run it

Run a short structured improv sequence — word association at speed, "yes, and" building, one-word-at-a-time storytelling — with the group standing, and with the explicit rule that the aim is not to be funny.

What makes this an evidence-informed activity rather than an icebreaker is the debrief. Improv's studied benefit is not team bonding, which is what it is usually sold as. It is a measured change in divergent thinking and in tolerance for uncertainty. Say that, and measure it.

Facilitator script

  1. 5 min · Set the anti-comedy rule
    The aim is not to be funny. If you are trying to be funny, you are editing, and editing is the thing this exercise exists to suppress. Speed and acceptance, nothing else.
  2. 10 min · Word association at speed
    Circle, one word each, no pauses. If you hesitate more than two seconds we start again. Nonsense is fine. Repetition is fine.
  3. 15 min · Yes, and
    In pairs, build a scenario. Every line begins by accepting what came before and adding one thing. You may not block, and you may not ask a question that hands the work back.
  4. 10 min · One word at a time
    Whole group, one story, one word each around the circle. Notice the moment you want to steer it somewhere — and let it go.
  5. 10 min · Debrief the real claim
    What this is evidenced for is divergent thinking and comfort with uncertainty. It is not evidenced as a team-building intervention, whatever the person who sold it to your company said. What actually changed in the last forty minutes?

The evidence

Felsman, Gunawardena & Seifert (2020), in Thinking Skills and Creativity and related work, compared short improvisation training against an active control activity and found gains in divergent thinking and in tolerance of uncertainty, with affect improvements alongside. The use of an active control rather than a do-nothing group puts this above most of the training literature Sio & Lortie-Forgues criticised.

The limits are the usual ones for this corner of the field: modest samples, mostly student participants, short follow-up, and little independent replication. The direction is consistent; the magnitude is not established.

Note what the studies do not show. The common corporate justification — improv builds trust and collaboration — is not what was measured. If you want the climate effect, activity 10 has far better evidence for it.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationSingle study, not independently replicated
OutcomeA validated creativity measure (AUT, RAT, insight)

Randomised comparison against an active control condition, with divergent-thinking and affect outcomes; small samples and limited independent replication.

Felsman, P., Gunawardena, S., & Seifert, C. M. (2020). Improv experience promotes divergent thinking, uncertainty tolerance, and affective well-being. Thinking Skills and Creativity, 35, 100632. doi:10.1016/j.tsc.2020.100632 Find this paper ↗

Did you knowImprov's core rule and Osborn's core rule are the same instruction — accept before you evaluate — arrived at independently by a theatre tradition and a psychology tradition. Activity 27 is the reason that convergence should not be taken as confirmation.
Watch outPhysical, performative exercises are not neutral. They exclude some disabled participants, they are much harder for non-native speakers at speed, and in a group with a real status gradient they can be humiliating rather than liberating. Offer a seated, written variant every time and mean it.
33Moderate

Activating the mood, not lifting it

Positive mood helps only if it is activating. Calm contentment does roughly nothing, and the distinction is the entire finding.

DRO
25 minAnyclimate
Full detail

How to run it

Most "get the room in a good mood" advice is wrong in a specific and correctable way. What the meta-analytic evidence supports is not positive mood but activating positive mood — high-arousal states such as elation or excitement. Low-arousal positive states like calm and relaxation show little effect on idea generation.

Run two short inductions and a divergent task after each, then have the room plot themselves on a two-by-two: pleasant/unpleasant against activated/deactivated. The two-by-two is what people take away.

Facilitator script

  1. 5 min · Draw the two-by-two
    Two axes, not one. Pleasant to unpleasant across, activated to deactivated up. Four quadrants. Most advice about mood and creativity only knows about the left-right axis, and that is why it under-performs.
  2. 8 min · Induction one, then task
    Short high-energy piece, then four minutes of divergent generation immediately. Do not discuss between them.
  3. 8 min · Induction two, then task
    Now something calm and pleasant, then the same four minutes on a different prompt.
  4. 4 min · Plot and count
    Mark yourself in a quadrant for each round and count your outputs. This is a demonstration, not an experiment — order effects alone could produce the difference. But the quadrant map is the durable part.

The evidence

Baas, De Dreu & Nijstad (2008), Psychological Bulletin, meta-analysed 25 years of mood-and-creativity research. The headline is the qualification: positive mood benefits creativity when it is activating; deactivating positive moods show little or no benefit. Activating negative moods such as anger can also raise output, though through a different route — persistence rather than flexibility — and at a cost worth thinking about before you engineer it.

This connects to the dual-pathway account (Nijstad, De Dreu, Rietzschel & Baas, 2010): creative output can be reached either by flexibility, moving between categories, or by persistence, going deep within them. Different states favour different routes, which is why a single "be positive" instruction under-performs.

Grade honestly. This is a large, well-conducted meta-analysis, and it is also mood-induction studies with lab divergent-thinking outcomes, mostly short-term, mostly students. And by Sio & Lortie-Forgues's own moderator analysis, motivational and affective routes underperform cognitive-skills training — so this is a supporting card, not a headline one.

Evidence profile

DesignMeta-analysis, or trial with an active control
ReplicationMeta-analytic or consistently replicated
OutcomeA validated creativity measure (AUT, RAT, insight)

Large meta-analysis of mood-induction experiments; outcome measures are predominantly lab divergent-thinking tasks.

Baas, M., De Dreu, C. K. W., & Nijstad, B. A. (2008). A meta-analysis of 25 years of mood-creativity research. Psychological Bulletin, 134(6), 779–806. doi:10.1037/a0012815 Find this paper ↗

Did you knowThe activation distinction quietly explains a common facilitation failure. Running a generation session immediately after a long, pleasant, low-energy lunch puts a room in the deactivated-pleasant quadrant — the one the evidence says does least.
Watch outDo not manufacture anger because activating negative moods also show effects. The persistence route is real and the organisational cost of deliberately provoking a team is not something a meta-analysis of lab inductions can tell you about.
34Contested

SCAMPER, audited

Probably the most-taught creativity technique on earth, and one of the least tested. Run it — then run the trace on why you had never questioned it.

DRO
40 minAnydiverge
Full detail

How to run it

SCAMPER — Substitute, Combine, Adapt, Modify, Put to another use, Eliminate, Reverse — was assembled by Bob Eberle from Osborn's checklist questions. It is a mnemonic over a list, not a research finding, and it has never been the subject of anything resembling the trials this site asks for elsewhere.

So run it twice. First as a technique, honestly, because it does generate ideas. Then as a source-tracing exercise: ask the room to find the study. Fifteen minutes with a search engine is usually enough for the point to land on its own.

Facilitator script

  1. 15 min · Run it straight
    Seven prompts, one pass each. Substitute what? Combine with what? Keep going through the list. Capture everything — this half is not a trick, the technique does produce output.
  2. 15 min · Now find the evidence
    Search for the controlled trial. Randomised, active control, creativity outcome. You have fifteen minutes. Tell me what you find and, more importantly, what the studies you do find actually compared against.
  3. 10 min · Separate the two questions
    Two different questions, and rooms collapse them constantly. Does this produce ideas? Almost certainly yes — any structured prompt beats staring at a wall. Does it produce more or better ideas than another structured prompt? Nobody has really tested it. Only the second question justifies teaching it as a method.

The evidence

SCAMPER's lineage runs from Osborn's Applied Imagination (1953) checklist questions to Eberle's mnemonic (1971). Neither is a controlled study. The published evaluation literature consists mainly of small quasi-experimental studies, frequently with school-age participants, frequently comparing a SCAMPER-based programme against no treatment, and almost never isolating SCAMPER from the surrounding instruction, facilitation and attention.

By the standards Sio & Lortie-Forgues apply — randomisation, an active control, a pretest — essentially none of it qualifies. This does not mean SCAMPER does not work. It means the claim has not been tested, which is a different statement, and the difference is what this whole site is about.

Compare activity 30. Design Heuristics prompts overlap substantially with SCAMPER prompts in content; the difference in grade is entirely a difference in who did the studies.

Evidence profile

DesignCorrelational or field-observational
ReplicationSingle study, not independently replicated
OutcomeAn adjacent cognitive outcome only

Widely used, weakly tested. Small quasi-experimental studies, mostly with children, mostly against no-treatment controls, rarely isolating SCAMPER from the training that surrounds it.

Eberle, B. (1971). SCAMPER: Games for Imagination Development. Buffalo, NY: D.O.K. Publishers. Derived from Osborn, A. F. (1953), Applied Imagination. Find this paper ↗

Did you knowThe most common defence is that SCAMPER is obviously useful, so studying it would be a waste. That argument was made for defer-judgment brainstorming for thirty years before Diehl and Stroebe measured it and found nominal groups outperforming real ones.
Watch outDo not use this card to tell a room their favourite tool is worthless. Contested means untested, not refuted, and a facilitator who overclaims in the sceptical direction has made exactly the error described in section 05, just pointing the other way.
35Contested

Six Thinking Hats, audited

The most commercially successful creativity framework in history has almost no controlled evidence — and one component that is well supported for entirely different reasons.

DRO
45 min4–20climate
Full detail

How to run it

De Bono's six hats assign the whole group a single mode at a time: white for information, red for feeling, black for caution, yellow for benefit, green for new ideas, blue for process. Run it on a real decision so people can feel what it does.

Then separate the framework from its one defensible mechanism. The instruction everyone does the same kind of thinking at the same time is a form of generate/evaluate separation, and that has independent support — see activity 13. The rest of the apparatus — six specific modes, these colours, this sequence — does not.

Facilitator script

  1. 5 min · Explain parallel thinking
    The rule is that we are all in the same mode at once. Nobody is the designated sceptic and nobody is the designated enthusiast. That single rule is the part worth keeping.
  2. 22 min · Run the sequence
    White first — only what we know, no interpretation. Then green, only new possibilities. Then yellow, then black, then red. Blue is mine. Roughly four minutes each, and stay in the mode.
  3. 8 min · Isolate the active ingredient
    Which part of the last twenty minutes did the work? Be specific. My claim is that it was the separation, not the colours — argue with me.
  4. 10 min · Check the evidence chain
    De Bono's method has sold in the millions and been adopted by governments. Search for the randomised trial with an active control. Then ask what commercial success is actually evidence of.

The evidence

De Bono's Six Thinking Hats (1985) is a practitioner framework. Its published evaluation base is dominated by case reports, organisational testimonials, satisfaction surveys and small studies without active controls. There is no body of randomised trials establishing that the six-mode structure outperforms another structured discussion protocol on creative output.

What can be defended is narrower. Separating generation from evaluation has good support (activity 13), including a mechanistic account in the network-switching literature in section 03. Reducing evaluation apprehension in groups has support (activities 07 and 10). The hats are one implementation of principles evidenced independently of the hats.

Teach it that way and you have kept everything usable while dropping the unsupported claim. That is a better outcome than either uncritical adoption or dismissal.

Evidence profile

DesignCorrelational or field-observational
ReplicationSingle study, not independently replicated
OutcomeAn adjacent cognitive outcome only

Very widely adopted, very thinly evidenced. Mostly case reports and satisfaction surveys; the one defensible component is separation of thinking modes, which is evidenced elsewhere.

de Bono, E. (1985). Six Thinking Hats. Boston: Little, Brown. Read against Nijstad & Stroebe (2006) on group ideation mechanisms. Find this paper ↗

Did you knowAdoption and evidence are close to uncorrelated in this field. The three most widely deployed creativity frameworks in corporate use all have weaker evidential support than several activities in this library that almost nobody has heard of.
Watch outBlack hat has an asymmetric cost. In a group with a status gradient, a scheduled criticism round licenses senior people to do what they were going to do anyway with official sanction. If you cannot run activity 10 first, consider running black hat in writing.
36Contested

The five-stage design thinking pipeline

The dominant framework in innovation education. Its components are variously well and poorly evidenced; the five-stage sequence itself is a teaching device that became a claim.

DRO
40 minAnyframe
Full detail

How to run it

Empathise, define, ideate, prototype, test. Put the diagram up and ask the room to grade each stage separately against the criteria in section 05 — what was measured, in whom, by whom.

The exercise works because the answers genuinely differ by stage. Prototyping and iterative testing have real support in product development research. Problem definition maps onto the problem-construction literature in activity 08, which is decently evidenced. The claim that these five stages in this order constitute a validated process is a different claim, and it has not been tested against any alternative ordering.

Facilitator script

  1. 5 min · Put the diagram up
    Five boxes, five arrows. Almost everyone in this room has seen it. Today we are grading it box by box rather than accepting or rejecting the whole thing.
  2. 18 min · Grade each stage
    For each stage: what is the primary evidence that this stage, as specified, improves outcomes? Not whether it feels right. What was measured, in whom.
  3. 9 min · Find the untested claim
    Now the sequence itself. Where is the study comparing this order against another order? If we cannot find one, the arrows are a pedagogical convenience that got promoted.
  4. 8 min · Keep what survives
    What would you still run tomorrow, and what would you stop calling evidence-based? That is a shorter list and a more honest one, and your students will trust it more.

The evidence

The five-stage formulation popularised by the Stanford d.school and IDEO is a teaching model, not a research finding. The surrounding literature is dominated by case descriptions and practitioner accounts; systematic reviews of design thinking's effectiveness repeatedly report heterogeneous definitions, weak measurement and few controlled comparisons.

Components fare differently. Iterative prototyping and early user testing have support in new product development research. Problem construction has support (activity 08). Empathy work as specified — short-form interviews producing personas — has notably little controlled evidence for its effect on creative output, whatever its other merits.

The interesting failure mode is the sampling frame from section 05. Much of the visible evidence for design thinking comes from consultancies and schools that teach it, which is not disqualifying but does mean the enthusiastic literature and the independent literature are not the same literature.

Evidence profile

DesignQualitative, case-based or theoretical only
ReplicationSingle study, not independently replicated
OutcomeAn adjacent cognitive outcome only

Framework-grade with a large case literature and a small, mixed outcome literature; the specific five-stage sequence has not been tested against alternative structures.

Brown, T. (2008). Design thinking. Harvard Business Review, 86(6), 84–92. Read against systematic reviews of design thinking effectiveness, e.g. Micheli et al. (2019), Journal of Product Innovation Management, 36(2), 124–148. doi:10.1111/jpim.12466 Find this paper ↗

Did you knowMicheli and colleagues' review found the design thinking literature could not agree on the construct's own attributes, let alone its effects — a definitional problem that makes a clean meta-analysis close to impossible.
Watch outThis card will annoy people, including people whose teaching depends on the framework. Run it as a grading exercise, not a takedown — the aim is a defensible version of design thinking, and there is one.
37Contested

TRIZ and the contradiction matrix

A method built from patent analysis rather than psychology, with a genuinely interesting derivation and almost no controlled testing.

DRO
50 min3–16diverge
Full detail

How to run it

TRIZ's core move is worth teaching regardless of its evidence grade: state the problem as a contradiction — improving A degrades B — rather than as a goal. Then look up the contradiction in the matrix and read off the inventive principles that historically resolved that pairing.

Run it on a real technical problem. The forced move from "make it better" to "these two parameters are in conflict" is where most of the value sits, and it survives even if the matrix does not.

Facilitator script

  1. 10 min · State the contradiction
    Not a goal. A conflict. Something of the form: if we increase this, that gets worse. If you cannot write it in that form, you have not analysed the problem yet.
  2. 8 min · Locate it in the matrix
    Find your two parameters, read the cell, note the principles it suggests. Do not evaluate them yet.
  3. 17 min · Apply the principles
    For each suggested principle, force an application. Segmentation — what would it mean to divide this? Prior action — what could be done in advance? Force it even where it seems absurd.
  4. 15 min · Audit the derivation
    Now the source question. Altshuller derived these by analysing patents. What kind of claim does that support, and what kind does it not? Specifically: a patent database tells you what was invented and granted. Does it tell you what works, or what gets filed?

The evidence

Altshuller derived TRIZ's 40 inventive principles and the contradiction matrix from analysis of large volumes of patents, identifying recurring resolution patterns. As a derivation this is more interesting than pure invention — it is at least grounded in a real corpus.

The evidential problem is what that corpus can support. Patent analysis is a study of what was filed and granted, which is not the same as what solved problems well, and it is silent on whether a person given the matrix outperforms a person given another structured method. The applied literature is dominated by industrial case studies and success reports, with very few controlled comparisons and almost no randomisation.

The contradiction formulation itself sits closer to defensible ground: it overlaps substantially with problem construction (activity 08) and with C–K's use of a knowledge constraint to open a concept (activity 26), both of which have better support.

Evidence profile

DesignQualitative, case-based or theoretical only
ReplicationSingle study, not independently replicated
OutcomeAn adjacent cognitive outcome only

Derived from large-scale patent analysis rather than experiment; adoption evidence is industrial case studies, with almost no controlled comparison against other structured methods.

Altshuller, G. S. (1984). Creativity as an Exact Science. Gordon & Breach. See also Ilevbare, Probert & Phaal (2013), Technovation, 33(2–3), 30–37, on adoption evidence. doi:10.1016/j.technovation.2012.11.003 Find this paper ↗

Did you knowTRIZ makes an explicit claim that most creativity methods avoid: that inventive solutions are not idiosyncratic but fall into a small number of recurring patterns. That is a testable claim. It has been argued about far more than it has been tested.
Watch outThe matrix is dated in a specific way — it was built from a mid-twentieth-century mechanical and chemical patent corpus, and it shows when applied to software, services or systems problems. Use the contradiction framing broadly; use the matrix where the problem looks like the corpus.
38Contested

The room-cue claims

Blue walls, high ceilings, coffee-shop noise. Three famous published effects, all of which have had a difficult decade. Also the card that corrects activity 14.

DRO
30 minAnystate
Full detail

How to run it

Do not run these as interventions. Run the trace, as with activity 19.

Hand out the three abstracts — blue versus red (Mehta & Zhu, 2009, Science), ceiling height (Meyers-Levy & Zhu, 2007), ambient noise at around 70 dB (Mehta, Zhu & Cheema, 2012). All three are published in strong outlets, all three are widely repeated in workplace design advice, and all three sit in the part of social and consumer psychology hit hardest by the replication crisis.

Ask the room how many of the three they had heard as settled fact before today.

Facilitator script

  1. 4 min · Poll first
    Hands up: who has heard that blue rooms help creativity? High ceilings? Moderate background noise? Keep your hands up so you can see how much of the room shares each belief.
  2. 12 min · Read the originals
    Three abstracts. For each: what was the sample, what was the outcome measure, and how large was the effect? Write the numbers down before you form a view.
  3. 8 min · Read what happened next
    Now the replication record. Note specifically whether the follow-ups were larger, preregistered, and by independent groups — and what happened to the effect when they were.
  4. 6 min · Reconcile with activity 14
    Here is a live contradiction inside this library. Activity 14 tells you to lower stimulation during generation. The 70-decibel finding says moderate noise helps. Both cannot be simply true. What would you have to know to decide?

The evidence

Colour. Mehta & Zhu (2009), Science, reported blue backgrounds favouring creative tasks and red favouring detail-oriented tasks. The effect has proved fragile: subsequent larger and preregistered attempts have found much smaller or null effects, and colour-cognition effects generally are among the more contested in the field.

Ceiling height. Meyers-Levy & Zhu (2007), Journal of Consumer Research, reported higher ceilings promoting abstract, relational processing. The construct-priming literature this belongs to has been badly damaged by replication failures across the board.

Ambient noise. Mehta, Zhu & Cheema (2012), Journal of Consumer Research, reported moderate ambient noise (around 70 dB) improving creative cognition relative to low or high noise. Replication attempts have been mixed, with several failing to recover the inverted-U pattern.

And the honest consequence for this library: activity 14 is not as secure as its moderate grade implies. The alpha-synchronisation and eye-closure evidence behind it is reasonable, but it points in the opposite direction to the noise finding, and this site should not carry one and hide the other. The current reading — that the two studies measure different things, internally-directed generation versus a broad creative-cognition composite — is a reconciliation, not a result.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationFailed a preregistered replication
OutcomeA validated creativity measure (AUT, RAT, insight)

Well-designed original experiments in good journals whose effects have weakened or failed under larger, preregistered replication attempts.

Mehta, R., & Zhu, R. (2009). Blue or red? Science, 323(5918), 1226–1229. Mehta, Zhu & Cheema (2012), Journal of Consumer Research, 39(4), 784–799. Meyers-Levy & Zhu (2007), Journal of Consumer Research, 34(2), 174–186. doi:10.1126/science.1169144 Find this paper ↗

Did you knowThese three findings have probably shaped more office refurbishments than the entire creativity training literature has shaped curricula. Effect on the built environment is not evidence of effect on people.
Watch outContested is not refuted, in either direction. None of these has been shown to be false; they have failed to hold up as robustly as their popularity assumes. Present them as unsettled, and resist the pleasure of debunking.
39Contested

Literally outside the box

The study where sitting outside a physical box improved creativity. A perfect teaching case: elegant design, good journal, and it did not survive.

DRO
25 minAnystate
Full detail

How to run it

Leung and colleagues (2012) reported that embodying metaphors improved creativity — including participants who literally sat outside a five-foot cardboard box performing better on creativity measures than those inside it.

Run this exactly as activity 19: as a trace, not a technique. It is arguably the best single teaching case in the library, because the original is genuinely clever and nothing about it looks careless.

Facilitator script

  1. 4 min · Tell it as a fact
    Here is a finding. People who sat outside a physical box scored higher on creativity tasks than people who sat inside one. Published in a leading journal. How plausible does that sound?
  2. 8 min · Read the original
    The design is good. Multiple experiments, several metaphors, sensible controls. Note what impresses you — that reaction is the thing we are examining.
  3. 8 min · Read the replication record
    Now the follow-ups. Larger samples, preregistration, independent labs. What happened to the effect?
  4. 5 min · Extract the rule
    Nobody here did anything wrong. The original was not sloppy and the replicators were not hostile. So what is the rule you take away? Mine is: a single elegant study is a hypothesis, and elegance is not evidence.

The evidence

Leung, Kim, Polman, Ong, Qiu, Goncalo & Sanchez-Burks (2012), Psychological Science, reported a set of embodied-cognition effects on creativity, including the literal box manipulation. The paper was widely covered and became a standard illustration in creativity teaching.

It belongs to the embodied-priming literature that the replication crisis affected most severely. Larger and preregistered replication efforts across this family of effects have generally failed to recover them at the reported magnitudes, and the field's confidence in short-manipulation priming of complex cognition has fallen substantially since 2012.

What makes it valuable teaching material is that the failure is not attributable to misconduct or incompetence. It is a well-executed study of an effect that appears not to be there, or to be far smaller than a small sample could detect. That is the ordinary case, and facilitators should expect it rather than treat it as scandal.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationFailed a preregistered replication
OutcomeA validated creativity measure (AUT, RAT, insight)

Original controlled experiments in a leading journal; subsequent larger and preregistered replications did not recover the effect.

Leung, A. K.-y., Kim, S., Polman, E., Ong, L. S., Qiu, L., Goncalo, J. A., & Sanchez-Burks, J. (2012). Embodied metaphors and creative acts. Psychological Science, 23(5), 502–509. doi:10.1177/0956797611429801 Find this paper ↗

Did you knowThe reason this one spread so fast is that it makes a pun literal. Findings that are also good jokes travel further than findings that are merely true, and that selection pressure operates on what reaches your slide deck.
Watch outDo not use this to teach that psychology is unreliable. The correct lesson is narrower and more useful: single studies of small priming effects on complex behaviour are weak evidence, whereas meta-analyses with bias correction and preregistered replications are the strong evidence — and this site is built on the second kind.
40Contested

The 0.075 finding

Mild intoxication improved performance on a convergent insight task. Included as a discussion card and a boundary case — emphatically not as an activity to run.

DRO
20 minAnystate
Full detail

How to run it

This is a discussion card. Do not run it. It is in the library because it comes up in every workshop that touches creativity and drinking, and a facilitator should be able to say precisely what the study did and did not show.

Jarosz, Colflesh & Wiley (2012) brought participants to roughly 0.075 blood alcohol content and tested them on the Remote Associates Test. The intoxicated group solved more problems, and faster, and were more likely to report the solution as arriving suddenly rather than analytically.

The teaching value is in the boundary conditions, which almost nobody who repeats this finding can state.

Facilitator script

  1. 4 min · State the finding precisely
    One specific blood alcohol level. One specific task — the Remote Associates Test, which measures convergent insight, not idea generation. One small sample. Say all three of those together or you have misreported it.
  2. 7 min · Find the boundary
    What would you predict for divergent generation? For sustained work? For the same people two hours later? The study speaks to none of those, and the popular version speaks to all of them.
  3. 9 min · Ask why it travelled
    This study is cited far more often in workshops than any of the meta-analyses on this site. Why? And what does that tell you about how findings are selected for transmission in your own field?

The evidence

Jarosz, Colflesh & Wiley (2012), Consciousness and Cognition, tested moderately intoxicated participants (around 0.075 BAC) against sober controls on the Remote Associates Test. The intoxicated group solved more items in less time and were more likely to characterise solutions as insight-like. The proposed mechanism is reduced executive control, broadening the search — the same mechanism proposed for the chronotype-mismatch effect in activity 15.

The limits are substantial. Small sample, a single task measuring convergent insight, no divergent measure, no assessment of real creative work, and no large replication. There is also a well-known asymmetry: the same reduced control that may help a three-word association task is very likely to hurt the sustained evaluative work that most creative production requires.

Included because it is a boundary case with a clear mechanism, and because being able to state a finding's limits precisely is the skill this library is actually trying to build.

Evidence profile

DesignControlled experiment, single laboratory
ReplicationSingle study, not independently replicated
OutcomeA validated creativity measure (AUT, RAT, insight)

Small controlled laboratory experiment with an intoxicated and sober comparison on a convergent insight measure; not replicated at scale, and not a runnable workshop activity.

Jarosz, A. F., Colflesh, G. J. H., & Wiley, J. (2012). Uncorking the muse: alcohol intoxication facilitates creative problem solving. Consciousness and Cognition, 21(1), 487–493. doi:10.1016/j.concog.2012.01.002 Find this paper ↗

Did you knowThe mechanism proposed here — reduced inhibitory control widening the search — is the same one Wieth and Zacks propose for the chronotype effect in activity 15. If you want the mechanism without the alcohol, run the divergent session at the wrong time of day.
Watch outNever run this as an activity. Facilitating a session involving alcohol is an occupational health matter, not a methodological one, and it excludes participants for medical, religious and personal reasons that are none of your business. It is on this site as a citation and a discussion, and that is all.
03 — Neuroscience

What the creative brain is actually doing

Neuroscience rarely tells a facilitator what to do. It does something more useful: it explains why the behavioural findings above take the shape they do — and it produces one genuinely counter-intuitive scheduling rule.

Creativity is network switching, not a "right brain"

Beaty and colleagues (PNAS, 2018) used connectome-based predictive modelling on fMRI from 163 participants doing a divergent thinking task. They identified a whole-brain network spanning three systems that normally work in opposition — the default mode network (spontaneous, associative thought), the executive control network (deliberate, goal-directed control), and the salience network (which arbitrates between them). Across four independent datasets, the strength of connectivity within that network predicted how original a person's ideas would be.

A 2025 multi-centre study in Communications Biology pushed this further: across 10 samples from five countries (N = 2,433), the number of dynamic switches between default and executive networks predicted creativity — but not general intelligence. Creativity looks like the capacity to alternate between generating and evaluating, fast and repeatedly.

Did you know The single most common facilitation error has a neural signature. Asking a room to "come up with ideas — but keep them realistic" demands simultaneous default and executive engagement of two systems that are anticorrelated at rest. Separating generation from evaluation isn't a nicety. It is working with the architecture rather than against it.
Close crop of hands writing independently on blank cards at a shared table
FindingWhat was measuredWhat it changes in the room
Alpha synchronisation EEG alpha power (8–13 Hz) rises reliably during idea generation, and rises more in more creative individuals. Fink & Benedek's review calls it among the most consistent findings in creativity neuroscience. It is read as internally-directed attention — the brain damping down external input. Generation is an inward task. Bright rooms, open laptops, and a facilitator talking over the silence all compete with the mechanism. Give people low-stimulation conditions and, if they'll tolerate it, permission to close their eyes. With one caveat this site owes you: a well-known published finding points the other way — moderate ambient noise improving creative cognition. Its replication record is poor, but it is not refuted. See activities 14 and 38.
The U-shaped alpha curve Alpha rises at the start of ideation, dips, then rises again just before a response is produced — most sharply in more creative individuals (Schwab et al., 2014). The mid-task dip is the felt experience of "running dry." It is a stage, not a stopping point — and it lands right where most brainstorms are cut short.
Chronotype mismatch Wieth & Zacks (2011): participants solved insight and analytic problems at their optimal or non-optimal time of day. Insight problem-solving was consistently better at the non-optimal time. Analytic problem-solving showed no consistent effect. Reduced inhibitory control is the proposed mechanism. Schedule the divergent workshop against the grain: late afternoon for larks, first thing for owls. Save the sharp hours for the analytic and decision work, which the time of day doesn't help anyway.
Sleep-onset (N1) Lacaux et al. (2021), Science Advances, N = 103: 15 seconds in N1 tripled the rate of discovering a hidden rule (83% vs 30%). The benefit disappeared entirely if participants dropped into deeper N2 sleep. The hypnagogic window is real and narrow. Taught as a personal technique rather than a workshop activity — see activity 06.
04 — Searching with AI

AI is a diverge-phase instrument

The most useful sentence in the current research on human–AI innovation search comes from a practitioner, not a theorist: AI can help you go broader more quickly. At some point you need to start to converge — and it is not so helpful on the converge piece.

Hassan & Nair (2026), in Research Policy, interviewed 41 people across 27 organisations and derived three relational dialectics — pairs where the human capability and the machine capability are opposed and mutually corrective. The framework maps cleanly onto the phase structure this site already uses, which is why it is worth teaching rather than just citing.

Human bringsAI bringsWhat it changes in practice
Human-centered search — tacit knowledge, trade secrets, relationships built over decades Boundary-spanning search — across languages, geographies, disciplines; and synthesis of internal knowledge the organisation has forgotten it owns Collapses the old local/distant distinction. In the age of big data, your own internal knowledge can be as inaccessible as anything outside the firm. See activity 21.
Contextualizing — relevance, situational meaning, forward-looking judgment about where a field is going De-contextualizing — no gut feeling, no prior dismissals, no domain assumptions filtering the search before it starts The clearest divergent/convergent split in the framework. Strip the brief of domain language before the machine sees it. See activity 20.
Judging — assessment, validation, and the wholly human work of persuading colleagues to back something Envisaging — visualising alternatives, and serendipity arising from outputs nobody asked for The weakest link in the framework, and the one needing the most structure. See activities 22 and 23.
The counterweight the framework almost ignores. Doshi & Hauser (2024), in Science Advances, ran an online experiment in which some writers received story ideas from GPT-4. AI-assisted stories were rated more creative, better written and more enjoyable — with the biggest gains among the least creative writers, effectively equalising outcomes. But those stories were measurably more similar to each other than the human-only stories. Individually better off; collectively, a narrower range of novel content. The authors call it a social dilemma. In the fifteen-page paper above, this finding appears as a single sentence in the final paragraph.

Which produces the practical question worth putting to any room: if the tool that widens your search is the same tool widening everyone else's, in the same direction, from the same training data — where does your distance advantage actually come from? The most likely answer is the half the machine is worst at. Your tacit knowledge, your relationships, and your judgment about what matters here.

05 — Source audit

Three ways evidence goes wrong in transmission

In every case below the original research is competent and properly published. The error enters afterwards — between the paper and the slide. These are the three failure modes worth being able to name, because once you can name them you stop needing anyone to vet claims for you.

Failure mode one — amplification

A real finding, in a small sample, gets restated as a universal protocol. Podcast neuroscience is where most professionals now meet this material, so a credible programme should engage with it directly rather than ignore it.

Andrew Huberman is a Stanford neurobiologist with a substantial publication record, and the Huberman Lab podcast has done more than any university to make protocol-level neuroscience legible to working professionals. He has also drawn sustained criticism from scientists and science journalists — Slate, Rolling Stone and New York among them — for extrapolating from animal studies, tiny samples and preliminary findings to confident, universal recommendations. Both things are true, and the useful move for an educator is neither adoption nor dismissal. It is checking the chain.

Protocol as popularisedWhat the primary study actually showsGrade
NSDR / Yoga Nidra "raises dopamine 65%" Kjaer et al. (2002), Cognitive Brain Research. Eight experienced yoga teachers, 11C-raclopride PET. Binding in the ventral striatum fell 7.9% during Yoga Nidra, which the authors calculate as a 65% rise in endogenous dopamine release. Real, careful, first-of-its-kind work — but n = 8, expert practitioners, and creativity was never measured. The dopamine-to-creativity step is an inference, not a finding. Moderate
90-minute ultradian focus blocks Traces to Kleitman's basic rest–activity cycle, proposed from sleep research and extended to waking cognition largely by analogy. Attention does fluctuate and breaks do help — but the specific 90-minute figure for creative knowledge work is a heuristic, not a measured optimum. Contested
Morning sunlight for cognitive performance Circadian entrainment via morning light exposure is genuinely well evidenced for sleep timing and alertness. There is no direct evidence linking it to creative output. Worth doing; not a creativity intervention. Strong (for sleep)
Deliberate cold exposure for focus and drive Small cold-water immersion studies have measured catecholamine changes, and the direction is consistent. The samples are tiny, the outcome measured is blood chemistry, and no study connects it to creative performance. It is also the protocol most likely to make a workshop cohort miserable. Contested
Panoramic / optic-flow vision to lower anxiety Mechanistically interesting and connected to Huberman's own visual-system research, but the applied claim outruns the published human evidence considerably. Treat as a hypothesis worth trying, not a protocol worth teaching. Contested

Failure mode two — the mechanism dies quietly

The headline result survives; the explanation underneath it is falsified years later, in a different literature, and nobody updates the slide. Activity 19 in the library is the worked example: thirty seconds of horizontal eye movements did raise two of five creativity sub-scores in a 2009 Brain and Cognition paper. The "it connects the two hemispheres" line that always accompanies it was the proposed mechanism, never a finding — and a preregistered adversarial collaboration in 2015 found strong Bayesian evidence against it, with a 2020 PLOS ONE study reaching the same conclusion.

The tell is grammatical. When a claim pairs a specific result with a confident causal story, check whether the second half was ever measured or only proposed.

Failure mode three — the sampling frame

The hardest to spot, because nothing about the study is wrong. Peer review passed, the method is appropriate, the analysis is careful. The issue is who was in the room.

Hassan & Nair (2026) — the human–AI search paper in section 04 — is a good, useful, open-access piece of work in a top journal, and this site teaches from it. It is also worth reading with three facts in view. Twenty-five of the twenty-seven participating organisations were recruited through the AI vendor's own customer network. Five of the forty-one interviews were with the vendor's own staff, including the CEO. And of the 64 secondary sources, roughly half are vendor-produced material — Amazon, Microsoft, Glean, Salesforce and Elastic blogs and customer stories — from which several of the paper's most quotable lines are drawn verbatim, including a "10 times faster" performance claim.

None of that makes the framework wrong. It is exploratory theory-building and does not claim to be anything else; the authors are notably candid, stating outright that they took "an overtly positive approach" and gave "little to no attention to the dark side." But it does mean the three dialectics are a usable framework, not a measured effect. No outcome was measured. Nobody tested whether search actually improved. And most of the people describing the tool's value had either bought it or sold it.

The question that catches all three. Not was this peer-reviewed? — all three cases were. Ask instead: what was actually measured, in whom, and who chose which people to ask? Teaching people to run that check themselves is more durable than teaching them any single activity in the library above.
06 — Ready to run

Three session blueprints

Cognitive content dominates deliberately — that is the moderator the evidence actually supports. Load any activity into the session builder to adapt these.

Where good ideas come from

90 minutes · any group size
  1. 0:00 Constraint demo — split room, blind rating
  2. 0:20 Problem reframing ladder, no solutions allowed — or the fixation demonstration (24) if the group has a house style to break
  3. 0:45 Walking pairs, 15 min outdoors
  4. 1:00 Silent brainwriting, 10 min enforced, then round-robin
  5. 1:20 Debrief: where in your list were your best ideas?

Generation and decision

Half day · 8–30 people
  1. Everything in the 90-minute design, plus
  2. The creativity bias module — delivered before convergence, ideally to decision-makers
  3. A real 20-minute incubation break with a low-demand task, between generation and selection
  4. Structured selection with a mandated high-novelty advance
  5. Schedule against chronotype: late afternoon for a morning-heavy team
  6. Optional source-audit hour — run activity 19, 34 or 39 as a trace exercise; it is the module that survives contact with a sceptical room

Innovation capability module

Six weeks · bachelor / EMBA
  1. Wk 1 The honest baseline + AUT pretest
  2. Wk 2 Problem construction, reframing, and the fixation demonstration (24)
  3. Wk 3 C–K mapping (26) and expansive partition (25); analogical transfer, distant-domain search
  4. Wk 4 Group process: production blocking, and structured dissent (27)
  5. Wk 5 Evaluation, the creativity bias, and creator-versus-manager forecasting (31)
  6. Wk 6 Live challenge, CAT-scored, AUT posttest
07 — Evaluation

If you claim a lift, measure it

Alternate Uses Task. List unusual uses for a common object. Score fluency, flexibility, originality, elaboration. Cheap, fast, well-validated — and what most of the studies above used.

Remote Associates Test. Three words, find the fourth that links them. Measures convergent insight. Used in the walking and nature studies.

Consensual Assessment Technique. Independent domain experts rate real products for creativity, blind, without a rubric. Highest ecological validity by a distance. Use this when assessing actual work.

Torrance Tests of Creative Thinking. The most-used instrument in the field's history, and the one to be most careful with. Its long-term predictive validity has been argued about for decades, scoring is proprietary and slow, and the norms are old. Know it because your literature is full of it — reach for it last.

Divergent Association Task. Name ten nouns as unrelated to each other as possible; originality is scored automatically from semantic distance. Four minutes, free, runs in a browser, validated against AUT and Bridge-the-Associative-Gap performance in Olson et al. (2021), PNAS. The most practical pre/post measure available to a working educator. doi:10.1073/pnas.2022340118

Automated semantic-distance scoring. Originality scoring of open-ended AUT responses by language model rather than by human raters (SemDis; Beaty & Johnson, 2021). It removes rater subjectivity and makes larger samples affordable — at the cost of measuring semantic distance, which is related to originality but is not the same construct. Report it alongside human scoring, not instead of it.

Design it the way the field failed to

  • Randomise at the individual level, not by class or team
  • Use an active control — a different training, not nothing. Only 35% of 169 studies did
  • Include a pretest, and a delayed posttest — only 23 of 169 did the latter
  • Power for d ≈ 0.3, not 0.6. Just 7% of published studies were adequately powered for the real effect
  • Preregister it. In fifty years and 169 studies, one person did. An OSF registration is free and takes an afternoon
Doing even three of these puts your evaluation in the top 11% of five decades of published creativity research.
Wall of clustered blank paper cards with a few coloured voting dots
Session 0 min