Critical Thinking Companion2026-09-06DRAFTsource .md12 sections

Critical Thinking Companion

Research brief and first product concept. Drafted 2026-09-06 from two deep research passes (theory and evidence; market and platforms), the identity, voice and offers files in _context/01_IDENTITY/, and what Palmer said on 6 September 2026. Every number in the market and evidence sections traces to a linked source. Prices and platform facts were read from the linked pages on 6 September 2026 by the research pass, and no vendor's live site or checkout was loaded in Palmer's own browser in this session; all of them get re-checked before launch.

What this is

A subscription on intelligencemeridian.com that helps a person keep and build the habit of thinking critically while AI takes over more of the thinking. Each week it gives the member a real thing to think about (a news item, an essay, a decision), a structured set of questions to work it with, a short piece of writing to do, and a small group of other people doing the same thing. The AI in it drafts questions and hints. It never writes the member's answer.

Palmer's own words, 6 September 2026: since leaving Columbia he has had a hard time finding people who engage with critical thinking. Most people let their emotions run their lives, which is unavoidable for a human being, and actually reading, writing and thinking critically about what is in front of your eyes is still extremely important. The product exists to make that engagement available on a weekly rhythm, with company.

The three decisions

These are the only things this document needs from Palmer. Everything below them is the reasoning.

1. Where the community lives. Recommendation: sell and gate on Squarespace, run the discussion on Skool. Squarespace member sites have no discussion, profiles or messaging (Kourses, August 2026), and MemberSpace confirms there is no native community feature (MemberSpace). Skool's Hobby plan is $9 a month with a 10 percent fee, which is the cheapest way to start, and its Pro plan is $99 a month at 2.9 percent once revenue passes about $1,300 a month (Kourses). Circle is the more structured option at $89 to $199 a month on annual billing (Ruzuku), which is the wrong bill for a product with zero members. Ship 30 for 30 runs on Skool (Ship 30 for 30). Rejected: Squarespace only (no discussion), Discord plus Patreon (14 to 16 percent of gross and weak for threaded critique, Ruzuku).

2. Price. Recommendation: $15 a month or $120 a year, a free weekly prompt carried in Meridian Passage as the top of the funnel, and a scholarship line for anyone who writes in. Comparable content plus community memberships without live coaching sit between $5 and $25 a month and $50 to $225 a year: Interintellect $22 or $225 (Interintellect), Farnam Street $25 or $149 (FS), Waking Up $19.99 or $129.99 (Waking Up), Ness Labs $49 a year (Circle). Annual prepay churns at roughly a third the rate of monthly (RetentionCheck), so the annual price is the one to push. The Self-Test at $39 stays below it as the gate it already is.

3. What the AI is allowed to do. Recommendation: the machine drafts question sets and hints against the week's material, Palmer edits them, and no AI writes or grades a member's response. This is the finding from the best designed study in the set. Bastani, Bastani and Sungu ran a randomized trial with nearly 1,000 students: unrestricted GPT-4 access raised practice scores 48 percent and then cut unassisted exam scores 17 percent below control, while a version that gave hints and withheld answers raised practice scores 127 percent and left exam scores level with control (Wharton summary, SSRN). Design decides whether the AI helps or harms. Say this on the sales page, because it is also the product's argument.

One more thing to hold in mind before answering. The weekly newsletter already claims 3 to 4 hours of production a week, and the Companion needs about the same. The concept below is built so the two share one weekly pass (the free prompt is a Passage section, and the member material is the same item worked deeper), because two separate weekly rhythms will not survive a full time job plus client work.

3 of 12. What critical thinking is, according to the people who defined itThere is no single agreed definition.

There is no single agreed definition. The frameworks disagree on whether critical thinking is a set of operations, a set of character traits, or something that only exists inside a domain you know well. A product that generates questions has to take a position on all three, because each implies a different kind of prompt.

Paul and Elder define it as the art of analysing and evaluating thinking with a view to improving it. Their model has eight elements of thought (purpose, question, information, concepts, assumptions, inferences, point of view, implications), nine intellectual standards (clarity, accuracy, precision, relevance, depth, breadth, logic, significance, fairness) and eight intellectual traits (humility, courage, empathy, autonomy, integrity, perseverance, confidence in reason, fair-mindedness) (Structural Learning). Their Miniature Guide pairs each standard with ready questions, for example "Could you give me an example?" for clarity and "Do I have any vested interest in this issue?" for fairness (Foundation for Critical Thinking, PDF). They also separate weak-sense thinking (using the tools to defend what you already believe) from strong-sense thinking (applying the standards to your own beliefs first). That distinction names the product's main failure mode: a member who gets better at arguing without getting better at thinking.

The 1990 Delphi Report (46 experts, six anonymous rounds) settled on "purposeful, self-regulatory judgment which results in interpretation, analysis, evaluation, and inference," with six core skills: interpretation, analysis, evaluation, inference, explanation and self-regulation, plus a list of dispositions such as inquisitiveness, open-mindedness, honesty about one's own biases and willingness to reconsider (Insight Assessment, full text). Self-regulation, meaning monitoring and correcting your own thinking, is named as a skill in its own right. That is why every week in the concept ends with a metacognitive question.

Robert Ennis: "reasonable reflective thinking that is focused on deciding what to believe or do" (HKU). His FRISCO checklist (Focus, Reasons, Inference, Situation, Clarity, Overview) is a procedure a person can run on any argument in a few minutes (Ennis summary).

Diane Halpern (1998) proposed four parts for teaching it so that it transfers: a dispositional component, instruction in specific skills, structure training so learners recognise the shape of a problem across contexts, and a metacognitive component (PubMed). Structure training is the strongest theoretical reason to apply the same question skeleton to news one week, an essay the next, and a personal decision the week after.

Daniel Willingham is the sharpest critic. In "Critical Thinking: Why Is It So Hard to Teach?" he argues that people see a problem's surface rather than its deep structure, and that without domain knowledge an instruction like "look at multiple perspectives" cannot be carried out (Reading Rockets). A 2023 review of systematic reviews reached the same place: instruction works better inside a subject than as a separate subject (Frontiers in Education). The answer the evidence supports is to make the material real and content rich, to pair it with a short reading that supplies the background, and to repeat the same structures across many domains until the shape becomes visible.

Skills versus dispositions. The 2023 review found far more studies treating critical thinking as a skill than as a disposition, and that the relationship between the two is unclear (Frontiers in Education). For the product this means the weekly practice has two jobs: skill drills, and cultivating the appetite for effortful thinking. The AI offloading research below suggests the appetite is the thing at risk.

4 of 12. What actually builds itThe evidence is modest and specific.

The evidence is modest and specific. The effects are real, they are small, and they come from a few named features.

Dialogue, real problems, mentoring. Abrami and colleagues analysed 341 effect sizes in 2015 and found an average effect of g = 0.30, with three instructional features standing out: opportunity for dialogue, exposure to authentic or situated problems, and mentoring (ERIC). Their 2008 analysis of 117 studies and 20,698 participants found g = 0.34 and concluded that critical thinking development "cannot be a matter of implicit expectation" (ERIC). This is the strongest backing for a product built on discussion, real material and some form of guided feedback, and the effect sizes are the reason the sales page should promise practice and habit rather than transformation.

Argument mapping. A meta-analysis of 26 studies reported effects around 0.7 for argument-mapping instruction, against roughly 0.11 for a semester of ordinary university attendance (van Gelder, PDF). Caveat: many of those studies were run by the technique's developers. One map a month is in the concept.

Consider the opposite. Lord, Lepper and Preston (1984) found that asking people how they would judge evidence if it pointed the other way reduced biased reasoning, and worked much better than telling them to be fair (SciSpace record). Prompts should force the opposite case rather than ask for fairness.

Spacing and retrieval. Cepeda and colleagues reviewed 839 assessments and found spaced practice beat massed practice in nearly every comparison, with the best gap growing as the retention horizon grows (Cepeda et al. 2006, PDF). Dunlosky rated practice testing and distributed practice as high utility, and rereading and highlighting as low (Dunlosky). A weekly cadence that brings back earlier question types at growing intervals fits.

Short writing with metacognitive prompts. Bangert-Drowns and colleagues found writing-to-learn effects were larger when the writing included metacognitive prompts, and smaller when the assignments were longer (Review of Educational Research). Short responses, with a "what was hardest, what would change your mind" question at the end.

Peer discussion. Smith and colleagues (Science, 2009) found that discussion improved performance on a second, structurally identical question even when nobody in the group had the first one right (ResearchGate record). Discussion itself produces understanding. This is the justification for the group, and for the order: answer alone first, then discuss, then revise.

Feedback in a wicked domain. Hogarth's distinction between kind learning environments (fast, accurate feedback) and wicked ones (delayed or misleading feedback) explains why experience alone does not build judgment about news or arguments (Epstein on Hogarth). The one practice with an objective score is forecasting: Tetlock's Good Judgment Project found a one hour training in probabilistic reasoning improved accuracy for at least a year, and that finding a base rate was the single most effective technique (AI Impacts, Good Judgment). One resolvable forecast a week gives the member a number that moves.

5 of 12. AI and cognitive offloading, 2023 to 2026This is the product's reason to exist, and the honest version is stronger than the alarming one.

This is the product's reason to exist, and the honest version is stronger than the alarming one.

Microsoft Research and Carnegie Mellon (CHI 2025). A survey of 319 knowledge workers with 936 examples of generative AI use found that higher confidence in the AI went with less critical thinking, and higher confidence in oneself went with more. Users described their thinking shifting toward verification, integration and stewardship (Microsoft Research). Self-reported and cross-sectional, so it documents perceived effort rather than measured skill.

Gerlich (2025). 666 UK participants, AI use correlated negatively with critical thinking (r = -0.68) and positively with cognitive offloading (r = 0.72); the 17 to 25 group used AI most and scored lowest (MDPI Societies). Correlational, convenience sampled, and the correlations are unusually large for a self-report instrument. A correction was issued in September 2025.

MIT Media Lab, "Your Brain on ChatGPT" (2025). 54 participants, EEG, essays written with an LLM, with search, or unaided. Connectivity was weakest in the LLM group, and LLM users struggled to quote their own essays (arXiv). A January 2026 commentary argues the study would need around 159 participants for power, that some comparisons rest on 2 to 4 essays, and that lower connectivity does not equal lower engagement (arXiv commentary). Widely quoted, still a preprint, cite with the caveats.

Bastani, Bastani and Sungu (2024). The randomized trial described in decision 3. Unrestricted access harmed later unassisted performance; the hint-only version did not (Wharton).

2026 work. A three wave survey of 589 students and early career workers separates dependent offloading (accepting the output as final) from autonomous offloading (using it as a starting point); dependent offloading predicted lower motivation and greater transfer of cognitive agency to the tool (Frontiers in Psychology, 2026). A Harvard physics trial with 194 students found a constrained AI tutor that withheld answers beat active learning classes by 0.73 to 1.3 standard deviations (Carl Hendrick's review).

Where they disagree. The surveys and the EEG study suggest broad erosion. The randomized trials say the effect depends almost entirely on whether the AI is allowed to do the thinking. The thread that runs through all of them: critical thinking declines when the person stops doing the effortful part and treats the output as final. The Companion is a weekly appointment with the effortful part.

6 of 12. What already existsNobody found in the research combines weekly prompts, generated question sets about a text or decision, and a community that critiques members' writing.

Nobody found in the research combines weekly prompts, generated question sets about a text or decision, and a community that critiques members' writing. The pieces exist separately.

Reasoning apps. Elevate sells 10 to 15 minute daily workouts at $9.99 a month or $39.99 a year, with a three games a day free tier (Nibble). Brilliant is $27.99 a month or $161.88 a year with an AI tutor inside courses and no community (E-Student). Clearer Thinking gives away more than 80 reasoning tools, runs a newsletter to over 400,000 subscribers, and makes money on coaching (Clearer Thinking). Socratic AI's slogan is "We don't give you answers, we teach you how to find them," free beta, no price (Socratic AI).

Writing and thinking communities. Ship 30 for 30 is $350 one time, self paced on Skool, built on one 250 word essay published daily for 30 days (Ship 30, review). Write of Passage charged $2,000 to $4,000 for five weeks, ran five person feedback pods matched by time zone, 15 to 20 person mentor groups, and a biweekly "Feedback Gym," and closed at the end of 2024 because AI changed Perell's model of writing education (Perell's version notes, Yespress). Interintellect caps salons near 25, requires active participation and runs its community on Discord (FAQ). Ness Labs runs weekly writing sessions and cohort "tiny experiments" at $49 a year, with an estimated 2,500 paying members and about 10 paid signups per newsletter issue (Latka). The Stoa, a philosophy community with loose structure, shows 82 paid Patreon members and a declining trend (Graphtreon). That last one is the cautionary case: without a bounded weekly action, a thinking community drifts.

AI study modes. OpenAI shipped study mode in July 2025 (OpenAI), Anthropic's learning mode reached all users in August 2025 (AlternativeTo), Google's Guided Learning followed the same month (TechCrunch). Every major lab now has a Socratic mode, mostly free. This validates the pedagogy and leaves the accountability and community layer open, which is the layer a solo operator can own.

What the successful ones share. A small bounded daily or weekly action. Public or semi-public output with feedback. Small groups (five in Write of Passage, 25 in Interintellect). An annual price under $250 unless there is live coaching.

Retention. Membership communities average about 5.8 percent monthly churn, low engagement is the top cancellation reason at 32 percent, and the first week is the biggest predictor of retention (RetentionCheck, Kourses). Cohorts with fixed dates complete at 85 to 96 percent against 3 to 6 percent for self paced courses (Wes Kao). A Dominican University study found 76 percent of people who wrote goals and sent weekly progress reports to a friend achieved them, against 43 percent who only thought about them (Dominican). These are editorial estimates and single studies; treat them as directional.

7 of 12. The product conceptA Brier score on their forecasts.

Name. Intelligence Meridian Critical Thinking Companion, shortened to the Companion. Working line: a weekly appointment with the part of thinking you are about to hand to a machine.

Fit with Intelligence Meridian. The company thesis is that a system has to be written down before automation accelerates whatever is already true about it. The Companion is that thesis applied to one person's mind: get your own reasoning legible to yourself before the tools make it invisible. The $39 Self-Test measures a person's legibility once. The Companion practises it weekly. The free prompt in Meridian Passage feeds both.

Who it is for. Three first members, in Palmer's order: knowledge workers who use AI daily and feel the muscle going, creatives and musicians from his own world, and students and recent grads who miss the seminar room. One product serves all three because the skeleton is the same and only the material changes. The material should rotate through their worlds.

The weekly ritual. Monday: the material lands (one item, a 300 to 600 word reading for background, the question set, one forecast). Wednesday: the member posts a response of 150 to 300 words, alone, before reading anyone else's. Thursday to Saturday: the pod of five reads and responds, each member required to steelman one other member's position. Sunday: the member writes three lines of revision (what changed, what did not, what would change it) and the forecast is logged. Palmer posts one short note on what the group saw. Total member time: 45 to 75 minutes a week. Total Palmer time: the same weekly pass that produces Passage, plus about one hour of reading and one note.

The material types, rotating. A news item. An essay or argument. A personal or professional decision the member brings. A forecast with a resolution date. Once a month, an argument map instead of prose.

The question engine. A generator that takes a text, decision or claim and produces a question set from the templates below, in a fixed order: purpose and question at issue, information and its gaps, the load-bearing assumption, the strongest opposite case, implications, then a self-regulation question. The machine drafts. Palmer edits and signs. The member sees the questions and a hint on request, never an answer. This is the piece to prototype first, because it is the piece that does not exist elsewhere and it is where the Intelligence Meridian mindset shows.

Tiers.

Tier Price What it is
Free $0 One prompt a week inside Meridian Passage. No group, no feedback.
Companion $15 a month or $120 a year Weekly material, question set, forecast, pod of five, Palmer's Sunday note.
Scholarship Free on request Same as Companion, for anyone who writes a paragraph saying why. Waking Up's precedent.
Consultation Hourly, existing offer For a member who wants one to one work on their own reasoning or their tools.

Community norms, in one paragraph. State your claim, your reasons, and what would change your mind. Criticise ideas, never people. Steelman before you disagree. Disclose any AI use in your response. Discussion of live party politics is out of scope unless the week's material is about it. Palmer moderates; repeated breaches end the membership. Drawn from Discourse's civility rules (Discourse) and LessWrong's epistemic guidelines (LessWrong).

What ships where. Squarespace: the sales page, checkout, the weekly material as gated pages, the Passage prompt. Skool: pods, responses, discussion, Sunday note. The question engine: a page on the site, or inside Skool as a link, once prototyped.

Progress a member can see. A Brier score on their forecasts. A count of steelmans another member accepted. A count of stated mind changes. Weeks completed. These track dispositions as well as skill, which the research says is the half most products ignore.

8 of 12. The first eight weeksEach week names the material type and the anchor question.

Each week names the material type and the anchor question. The full question set comes from the engine.

  1. A news item. What is the question this article is actually answering, and what question is it letting you assume it answered?
  2. An essay. What is the one assumption this argument cannot survive without, and how would you check it?
  3. A decision you are facing. Write the case for the option you are leaning against, well enough that a friend who holds it would sign it.
  4. A forecast. Pick a claim in the news that resolves within 30 days. Find the base rate. State a probability. Say what would move it ten points.
  5. An argument map. Take one op-ed and lay it out as boxes and arrows. Where does it thin out.
  6. A piece of AI output. Take something a model wrote for you this week. What did you verify, what did you accept, and what could you not reconstruct in your own words.
  7. A disagreement. Bring one you had recently. What would have to be true for the other person to be right.
  8. Revisit week. Return to week 1's article with what you know now. Score your week 4 forecast. Three lines on what changed.
9 of 12. The question templatesThirty templates the engine draws from, grouped by source.

Thirty templates the engine draws from, grouped by source. The full list with citations sits in the research pass, and these are the ones to build first.

Elements of thought (Paul and Elder): What is the author's purpose, and what would count as success. What is the precise question at issue, in one sentence. What information is offered, what is missing, and how would I check it. What concept does this depend on, and is it used consistently. What must be assumed for this to follow, and which assumption is weakest. What other conclusions fit the same evidence. From whose point of view is this written, and who would tell it differently. If this is true, what follows, and what follows from that.

Intellectual standards: Could I explain this to someone unfamiliar with it, with an example. How could I find out whether this is accurate. What complexities make this harder than it looks. Is this the most important issue here, or the most vivid one. Do I have a vested interest in this being true.

Socratic and Ennis: Why do you say that, and how do you know. What would be an example, and what would be a counterexample. What could we assume instead. How credible is this source, and what would make it more or less so. Why is this question worth asking.

Debiasing: If the evidence had pointed the other way, would I have accepted it as readily. Write the strongest version of the view you disagree with; would its holder sign it. What would have to be true for the other side to be right. Am I searching for support for one side only, and what did I skip.

Forecasting: What is the reference class and the base rate. State a probability; what evidence moves it ten points either way. Break the claim into sub-questions; which one carries the weight. What did I predict last month, what happened, and what did I get wrong.

Self-regulation: What single sentence would change my mind. Where am I least confident, and why. What part of this did I hand to a tool, and what did I verify myself. Explain the reasoning without looking at the source; what could I not reconstruct.

10 of 12. Risks, stated plainlyThe evidence for building critical thinking is modest (g around 0.3) and the alarming AI findings are mostly correlational or preliminary.

The evidence for building critical thinking is modest (g around 0.3) and the alarming AI findings are mostly correlational or preliminary. The sales page promises a weekly practice with company, and nothing more. Willingham's objection stands: a person cannot think critically about a domain they know nothing about, so the reading that accompanies each item is load bearing, and the material has to rotate through the members' own worlds. Time is the real constraint: the newsletter and the Companion have to be one weekly pass or one of them dies. The roles cap is full; this sits inside role 1 and displaces nothing, as long as it stays at the size described. The Stoa is what this becomes without a bounded weekly action and a fixed cadence. And the first week decides retention, so onboarding (pod assignment, first prompt, first response) has to be built before the first sale, the way the $39 product's checkout should have been tested before it went live.

11 of 12. Next steps, if the three decisions are yesPrototype the question engine as a single page: paste a text, get the ordered question set, with Palmer's edit pass before anything is shown to a member.

Prototype the question engine as a single page: paste a text, get the ordered question set, with Palmer's edit pass before anything is shown to a member. Write the sales page against voice.md. Draft the norms document. Set up Skool Hobby and one pod of five drawn from people Palmer already knows in the three audiences, and run the eight weeks free as shadow testing, which is consistent with the standing openness to free work while the business is new. Sell the annual after week eight to that pod first. Record the three decisions in _operating/registers/decisions.md when they are made.

12 of 12. SourcesTheory and evidence: Paul-Elder framework, Miniature Guide, Delphi Report, Facione 1990 full text, Ennis definition, Ennis FRISCO, Halpern 1998, Willingham, 2023 review of reviews, Abrami 2008, Abrami 2015, van Gelder argument mapping, Lord, Lepper and Preston 1984, Cepeda 2006, Dunlosky, Bangert-Drowns 2004, Smith 2009, Hogarth via Epstein, Good Judgment Project evidence, Tetlock's commandments.

Theory and evidence: Paul-Elder framework, Miniature Guide, Delphi Report, Facione 1990 full text, Ennis definition, Ennis FRISCO, Halpern 1998, Willingham, 2023 review of reviews, Abrami 2008, Abrami 2015, van Gelder argument mapping, Lord, Lepper and Preston 1984, Cepeda 2006, Dunlosky, Bangert-Drowns 2004, Smith 2009, Hogarth via Epstein, Good Judgment Project evidence, Tetlock's commandments.

AI and offloading: Lee et al., Microsoft Research, Gerlich 2025, Kosmyna et al. 2025, commentary on Kosmyna, Bastani et al., Wharton, Bastani SSRN, Zhu et al. 2026, Hendrick review.

Market and platforms: Elevate review, Brilliant review, Clearer Thinking, Socratic AI, Ship 30 for 30, Ship 30 review, Write of Passage version notes, Perell profile, Interintellect community, Interintellect FAQ, Ness Labs on Circle, Ness Labs revenue estimate, Farnam Street membership, Waking Up pricing, The Stoa on Graphtreon, Squarespace member areas review, MemberSpace on community, Skool pricing, Circle pricing, Patreon pricing, churn benchmarks, member retention, Wes Kao on cohorts, Dominican goals study, Discourse rules, LessWrong guidelines, OpenAI study mode, Claude learning mode, Gemini Guided Learning.

Sections in this document