Chapter 11
The lens, not the cage
In a darkened lab, a volunteer watches three blue squares appear on a screen, one above and two below. One of the bottom squares is the same blue as the top one, and the task is to choose it as quickly as possible. The instructions are simple until eight digits must be silently rehearsed at the same time. The digits are less considerate.
English has one basic color word for the whole region: blue. Russian ordinarily divides it, with goluboy for lighter blues and siniy for darker ones. Jonathan Winawer and colleagues brought speakers of both languages into the lab and timed them on the square task (Winawer et al. 2007).
In the original study, Russian speakers were faster on the difficult, near-color trials when the candidate squares crossed their own goluboy/siniy boundary. The advantage was 124 milliseconds without a second task and 109 milliseconds while they remembered a spatial pattern. English speakers showed no corresponding boundary advantage.
Then the researchers occupied the volunteers' inner speech. During the verbal condition, each person silently rehearsed an eight-digit string while judging the squares. The Russian category advantage disappeared. In those data, language was helping with a discrimination in the moment, and the help stopped when the verbal machinery was busy elsewhere.
That would make a wonderfully neat opening if the evidence had agreed to stop there. A preregistered reassessment using the same kind of discrimination task found no response-time advantage at the Russian light-blue/dark-blue boundary, although it did find one at the blue/green boundary (Martinovic, Paramei and MacInnes 2020). The later experiment was not exact, partly because calibrated monitors would not reproduce every original color, but it directly challenges the result on which the neat story rests.
The opening therefore has to hold both findings: one study found an effect with a verbal off switch, and a later study did not recover it. This is less satisfying than a parable and more useful than one.
What a real effect looks like
Phi was built on the idea that a language's structure matters to the people who use it, so a book about Phi owes its reader the evidence for that idea at its actual size. The strong claim is unsupported. Weaker effects differ by domain and task, and some admired examples have failed when tested again. This chapter keeps the accounts, including the crossed-out entries.
The strong claim says that language determines thought, that speakers of different languages inhabit mutually unreachable mental worlds, or that an unnamed concept cannot be thought. Reviews of linguistic relativity do not support that position (Wolff and Holmes 2011). A recent neuroscience review likewise assembles evidence that complex thought, including symbolic thought, does not require an intact language system, even though language can assist particular tasks and transmit cultural knowledge (Fedorenko, Piantadosi and Gibson 2024). That leaves room for a lens, but none for a locked room.
The weaker research program asks whether labels, grammatical habits, or familiar ways of speaking can alter attention, memory, speeded judgment, and the representations assembled for a task. Terry Regier and Paul Kay called the moderate position in color research "Whorf was half right" (Regier and Kay 2009), but even that convenient fraction covers several mechanisms and a disputed record. Number words can help preserve exact quantity; spatial language can travel with habits of orientation; a label can be recruited while colors are compared. The answers vary, and none licenses the conclusion that grammar quietly installs a worldview.
Two habits keep the account honest. Magnitudes stay attached to claims, so 124 milliseconds does not become a philosophy of blue, and corrections stay beside the findings they correct. The Russian result has its 2020 challenge here. So does the once-celebrated claim that color-category effects occur mainly in one visual field, which did not survive a series of ten experiments with 230 observers (Witzel and Gegenfurtner 2011). Replication belongs in the story rather than in an embarrassed appendix.
The sun on the table
In Pormpuraaw, an Aboriginal community on the west coast of Cape York, Lera Boroditsky and Alice Gaby gave people shuffled cards showing temporal progressions, such as a man aging or a banana being eaten, and asked them to put the cards in order (Boroditsky and Gaby 2010). The Kuuk Thaayorre-speaking participants tended to arrange the sequences from east to west. Facing south, they laid the cards left to right; facing north, right to left; facing east, toward the body; facing west, away from it. Every participant in the Stanford comparison group worked from left to right.
Kuuk Thaayorre commonly uses cardinal directions where English uses body-centered terms such as left and behind. Speakers must stay oriented well enough for those descriptions to work. The card task never asked about direction, yet the absolute frame appeared in the arrangement of time. Fourteen Pormpuraaw participants produced a striking and stable pattern, but fourteen people are still fourteen people, and the authors said as much.
The experiment shows covariation: a community that speaks about space in absolute directions also arranged temporal sequences on an absolute axis. It cannot decide whether language built the habit, reinforced it, or accompanies a broader cultural practice that built both.
That question produced a memorable exchange over rotated tables. Peggy Li and Lila Gleitman showed that landmarks could lead many English-speaking participants to use absolute responses, and that small matched cues could steer the strategy still further (Li and Gleitman 2002). Some of the fate of linguistic relativity therefore rested, for a time, on the placement of toy ducks. Experimental psychology is often more hands-on than its abstracts suggest.
Stephen Levinson and colleagues replied that Li and Gleitman had blurred distinct spatial frames and misread the original comparison; their own matched experiments continued to support covariance between habitual language and spatial reasoning (Levinson et al. 2002). The exchange confirmed durable differences in favored spatial frames while leaving their cause disputed. It also made the influence of task conditions impossible to ignore, which is a sounder result than either side's title promised.
Number is a technology
The Pirahã, an Indigenous people of the Brazilian Amazon, speak a language with no exact number words. Peter Gordon reported that Pirahã adults grew inexact beyond roughly three on several matching tasks (Gordon 2004). Popular retellings soon compressed this into a determinist slogan about number words and exact thought, well beyond what the data warranted.
Michael Frank and colleagues first tested the supposed number words and found that their use varied with context rather than exact cardinality. A form used for the smallest amount could still be used when six objects remained. The language did not have loose words for exact one and two; the tested forms expressed relative quantities instead (Frank et al. 2008).
The matching results drew a more useful boundary. Participants made exact matches with larger sets when they could pair the objects directly, item by item. Performance fell when exact cardinality had to survive a spatial rearrangement, concealment, a delay, or a change from one kind of object to another. They could enact direct correspondence but had no reusable symbol for carrying a set's exact size into the next situation.
Frank and colleagues therefore called number words a cognitive technology. The phrase does not mean that words create quantity. It means that a stable numeral can preserve exact cardinality when the objects themselves no longer keep their places. Spools and balloons did a modest amount of philosophical work in this study, as laboratory equipment sometimes does.
Phi's numerals are a technology chosen in the open. Three digits, mu ta wi, zero, one, two, combine with named powers of three, and the structure remains audible:
mia wi phoi ta shao torua phelu. 1SG two nine-group one three-group year hold. (I hold twenty-one years.)
Two nines and a three make twenty-one. Exact Phi numerals run through 242 and stop. Larger exact values, fractions, formulas, and measured units remain outside the Phi passage in their source notation. A speaker can state an age or divide a harvest within the range. A dosage of 325 milligrams stays beside the Phi account rather than being bent into it. The boundary is a design choice about which counting tools the language keeps inside itself, not a claim about what its speakers can understand.
The stories everyone remembers
Some claims about language and thought have acquired a second life as anecdotes. They travel well because they are compact and surprising, and because someone has kindly left their methodological luggage at the station. This section goes back for it.
The popular Hopi story says that Hopi has no time. Whorf's actual claims were more specific and have been interpreted in competing ways, but he did write that Hopi lacked forms referring to time as a smooth continuum familiar from European languages. Ekkehart Malotki answered with 677 pages of temporal expression and natural discourse (Malotki 1983), a notably patient way to disagree. His study decisively refuted the literal timeless-language story; scholars still debate whether it also disposed of Whorf's narrower argument about how Hopi organizes temporal reference. Hannah McElgunn's archival work adds an important human detail: Whorf developed much of his Hopi analysis through seven years of meetings with Ernest Naquayouma, a Hopi speaker then living in New York (McElgunn 2024). This was neither solitary genius nor a grammar written entirely from within the community.
The fixed number of Inuit words for snow is an academic urban legend. Franz Boas discussed four lexically distinct forms while illustrating how languages build words; later retellings promoted the count to fifty and then to hundreds. Geoffrey Pullum traced the inflation with deserved impatience (Pullum 1991). The correction also needs care. Inuit and Yupik languages are polysynthetic, so deciding what counts as one word is not simple, and individual languages have detailed vocabularies grounded in local knowledge of snow and ice (Krupnik and Müller-Wille 2010). What fails is the enormous fixed count, not the specialized knowledge.
The Mandarin-time story has also grown smaller under examination. Lera Boroditsky reported that Mandarin's vertical metaphors encouraged vertical representations of time (Boroditsky 2001). Three published papers challenged or failed to reproduce important parts of the result, and Jenn-Yeu Chen found that horizontal temporal metaphors were more frequent than vertical ones in Mandarin (Chen 2007; January and Kako 2007; Tse and Altarriba 2008). Later work using different, nonverbal tasks found a graded group difference: Mandarin speakers were more likely than English speakers to place earlier events above later ones, although both groups readily used a horizontal axis (Boroditsky, Fuhrman and McCormick 2011). A narrower claim survived after the dramatic one did not.
Then there are the gendered bridges. A famous report says German speakers described a grammatically feminine bridge as elegant and slender while Spanish speakers, for whom the noun is masculine, called it strong and sturdy. The result appeared briefly in an edited book chapter without enough procedural detail for a faithful replication (Boroditsky, Schmidt and Phillips 2003). Two later experiments failed to reproduce it (Mickan, Schiefke and Stefanowitsch 2014). A systematic review of 43 studies and 5,895 participants found that evidence for grammatical-gender effects depended heavily on language structure, task, and context, with several plausible explanations besides a lasting change to concepts (Samuel, Cole and Eacott 2019). The bridge remains useful as a warning about memorable examples, but it is poor building material for a general theory.
The savings claim remains livelier and less tidy. M. Keith Chen reported that speakers of languages with weak grammatical separation between present and future saved more and showed other future-oriented behaviors (Chen 2013). Controls for language history and geography weakened the association; under the strictest tests with individual-level data, it was not significant, although it remained in several other models (Roberts, Winters and Chen 2015). Later experiments have pointed both ways. A bilingual study reported more temporal discounting when the same participants worked in the language that marked future time more distinctly (Ayres, Kricheli Katz and Regev 2023). An English-Dutch study published in 2025 found that greater future-tense use predicted less discounting in English, the opposite of the original causal account (Robertson et al. 2025). The cross-country correlation cannot tell whether grammar causes behavior, shares a history with it, or travels beside another cause. The experimental question is still open.
Not every effect disappears, but fame selects for a cleaner story than research usually supplies. Milliseconds and source-memory differences can survive, as can habits of spatial layout. Worldviews are considerably harder to fit into the data.
This matters to Phi more than to most languages in this literature, because Phi is the one that was built on purpose. A natural language's speakers inherited their grammar and owe it no defense. A designed language arrives with a designer's hopes attached, and every story in this section is a hope of that kind that grew in the retelling until the data could no longer hold it up. The eight stations of the sun, the unpriced day, the syllable of admitted assumption: each of them is a design intention with a plain description available, and the plain description is what the rest of this book has been giving them.
The strongest counsel
The skeptical case deserves to be heard in its strongest form.
Steven Pinker begins with the observation that thought cannot simply be inner speech. Preverbal infants reason, people with severe aphasia can reason, and writers routinely possess a thought before finding a sentence fit to carry it (Pinker 1994, ch. 3). His sharper warning concerns argumentative method. Accounts of linguistic relativity can slide from a plain fact, languages differ in what speakers routinely encode, to a dramatic conclusion about different mental worlds. The certainty belongs to the first proposition and cannot simply be lent to the second.
John McWhorter adds two cautions (McWhorter 2014, chs. 1 and 2). Small experimental effects are often retold as worldviews, with milliseconds promoted to metaphysics. Exoticism also distorts judgment. Bulgarian marks information source grammatically and Greek does not, yet popular accounts rarely imagine a uniquely skeptical Bulgarian mind. McWhorter's point is ethical as well as empirical: flattering claims about how a distant people must think can still be condescending.
Evidentiality is the right place to test that warning because it is Phi's own territory. In Turkish past-tense clauses, one form marks firsthand knowledge while another marks non-firsthand information, which may come from hearsay or inference. Sümeyra Tosun and colleagues presented Turkish and English speakers with written assertions. Turkish adults later remembered firsthand-marked assertions better than non-firsthand ones; English speakers showed no corresponding source-type difference, and late Turkish-English bilinguals showed a related asymmetry when tested in English (Tosun, Vaid and Geraci 2013).
The task matters. Those participants were remembering linguistically framed reports and the form in which they had read them. In a later nonlinguistic event-memory study, Turkish and English speakers made equivalent errors when distinguishing events they had seen from events they had inferred (Ünal et al. 2016). Korean children, despite learning evidential morphology, developed nonlinguistic source reasoning on much the same timetable as English-speaking children (Papafragou et al. 2007). A broader review argues that evidential systems build on source concepts available across languages rather than creating those concepts (Ünal and Papafragou 2018).
The result is task-bound. Tosun and colleagues found a memory asymmetry for marked assertions, while the later work found no general advantage in nonlinguistic source monitoring. Grammar gave people practice with a distinction, not an epistemic virtue.
Phi cannot simply borrow even the first Turkish result. Its evidential particles are optional, and an unmarked sentence claims no source. The grammar offers a place where a speaker may disclose how they know, but it does not force the disclosure. Phi has four such markers: hi for witnessed knowledge, ke for inference, ti for report, and ho for assumption. An evidential is one syllable before the verb:
shia to ke wepu. 3SG PST INFER go. (She left [I infer from evidence].)
Past marker, inference marker, verb: she left, and the speaker admits that the leaving was worked out from evidence rather than watched. The boatman's narrator in this book's opening pages does the same thing. His expectation of payment arrives as to ho remo, past tense followed by admitted assumption. The sentence exposes the weakness of his knowledge before the story proves him wrong. Whether years of choosing such markers would alter a speaker's later habits is unknown. Phi offers the practice; it has not yet supplied the experiment.
An untested instrument
Constructed languages have been proposed as unusually clean tests of linguistic relativity for seventy years. The record of actually running such a test is short.
James Cooke Brown began Loglan in 1955 and presented it in Scientific American as a logical language that could help test the hypothesis (Brown 1960). The community timeline records an unsuccessful National Science Foundation proposal in 1977 and an apprenticeship that produced extended conversation, with pauses for dictionary work included at no extra charge (Lojban community n.d.). Disputes over control and trademark later helped produce Lojban. No published controlled test of Loglan or Lojban cognition was located for this audit. That is a report about the record searched, not proof that no informal attempt ever occurred.
Láadan came with its own deadline. Suzette Haden Elgin created the language in 1982 while writing Native Tongue and proposed that women would either adopt it or construct a better language in response. Ten years later she judged that hypothesis false (Elgin 1999). Her verdict did not mean that nobody used Láadan; a small community preserved and extended it. It meant that neither response occurred at the scale her proposed experiment required, so the cognitive hypotheses remained anecdotal.
Toki Pona presents a different case. It has a living international community, not a vanished test population. The self-selected 2024 census received 1,997 responses; 1,596 respondents said they knew Toki Pona, and 53 percent reported at least conversational ability (Toki Pona census 2024). Those numbers are not a population estimate, but they make the existence of an active community difficult to misplace. Learners often describe simpler thinking or gentler attention, yet no controlled cognition study was located for this audit (Lang 2014).
The obstacles therefore do not rhyme quite as neatly as one might wish. Loglan did not assemble and fund the proposed experiment, while Láadan did not receive the broad adoption its test required. Toki Pona achieved community but has not turned testimony into controlled evidence. A constructed-language study needs speakers and a question precise enough to fail, followed by researchers willing to ask it. Obtaining the whole arrangement at once has proved difficult.
Phi occupies an earlier position still. The most it may claim today is close to the position chapter 3 built on, Dan Slobin's "thinking for speaking": a language's requirements shape what speakers attend to while preparing an utterance (Slobin 1996). Phi places relations before content, puts a question where its answer would stand, and offers a slot for information source. These choices plainly organize Phi sentences, and anything past that is a hypothesis, which is where the evidence in this chapter leaves the whole question and where the protocol had already put it: Phi's structures are "design intentions to test in use, not guarantees about what the language does to a speaker."
References
Ayres, Ian, Tamar Kricheli Katz, and Tali Regev. 2023. Languages and future-oriented economic behavior—Experimental evidence for causal effects. Proceedings of the National Academy of Sciences 120(7): e2208871120. https://doi.org/10.1073/pnas.2208871120. Correction published 2024, 121(8): e2400771121. https://doi.org/10.1073/pnas.2400771121.
Boroditsky, Lera. 2001. Does language shape thought? Mandarin and English speakers' conceptions of time. Cognitive Psychology 43: 1-22. https://doi.org/10.1006/cogp.2001.0748.
Boroditsky, Lera, and Alice Gaby. 2010. Remembrances of times east: absolute spatial representations of time in an Australian Aboriginal community. Psychological Science 21: 1635-1639. https://doi.org/10.1177/0956797610386621.
Boroditsky, Lera, Orly Fuhrman, and Kelly McCormick. 2011. Do English and Mandarin speakers think about time differently? Cognition 118: 123-129. https://doi.org/10.1016/j.cognition.2010.09.010.
Boroditsky, Lera, Lauren A. Schmidt, and Webb Phillips. 2003. Sex, syntax, and semantics. In Language in Mind: Advances in the Study of Language and Thought, edited by Dedre Gentner and Susan Goldin-Meadow, 61-79. Cambridge, MA: MIT Press. https://doi.org/10.7551/mitpress/4117.003.0010.
Brown, James Cooke. 1960. Loglan. Scientific American 202(6): 53-63. https://doi.org/10.1038/scientificamerican0660-53.
Chen, Jenn-Yeu. 2007. Do Chinese and English speakers think about time differently? Failure of replicating Boroditsky (2001). Cognition 104: 427-436. https://doi.org/10.1016/j.cognition.2006.09.012.
Chen, M. Keith. 2013. The effect of language on economic behavior: evidence from savings rates, health behaviors, and retirement assets. American Economic Review 103(2): 690-731. https://doi.org/10.1257/aer.103.2.690.
Elgin, Suzette Haden. 1999. Láadan, the Constructed Language in Native Tongue. Essay first published on the author's SFWA member site, updated 2 April 1999; now hosted, under a slightly expanded title, at https://laadanlanguage.com/articles/articles-by-suzette/laadan-constructed-language/.
Fedorenko, Evelina, Steven T. Piantadosi, and Edward A. F. Gibson. 2024. Language is primarily a tool for communication rather than thought. Nature 630: 575-586. https://doi.org/10.1038/s41586-024-07522-w.
Frank, Michael C., Daniel L. Everett, Evelina Fedorenko, and Edward Gibson. 2008. Number as a cognitive technology: evidence from Pirahã language and cognition. Cognition 108: 819-824. https://doi.org/10.1016/j.cognition.2008.04.007.
Gordon, Peter. 2004. Numerical cognition without words: evidence from Amazonia. Science 306: 496-499. https://doi.org/10.1126/science.1094492.
January, David, and Edward Kako. 2007. Re-evaluating evidence for linguistic relativity: reply to Boroditsky (2001). Cognition 104: 417-426. https://doi.org/10.1016/j.cognition.2006.07.008.
Krupnik, Igor, and Ludger Müller-Wille. 2010. Franz Boas and Inuktitut terminology for ice and snow: from the emergence of the field to the "Great Eskimo Vocabulary Hoax." In SIKU: Knowing Our Ice, edited by Igor Krupnik, Claudio Aporta, Shari Gearheard, Gita J. Laidler, and Lene Kielsen Holm, 377-400. Dordrecht: Springer. https://doi.org/10.1007/978-90-481-8587-0_16.
Lang, Sonja. 2014. Toki Pona: The Language of Good. Self-published; official site https://tokipona.org/.
Levinson, Stephen C., Sotaro Kita, Daniel B. M. Haun, and Björn H. Rasch. 2002. Returning the tables: language affects spatial reasoning. Cognition 84: 155-188. https://doi.org/10.1016/S0010-0277(02)00045-8.
Li, Peggy, and Lila Gleitman. 2002. Turning the tables: language and spatial reasoning. Cognition 83: 265-294. https://doi.org/10.1016/S0010-0277(02)00009-4.
Lojban community. n.d. Lojban timeline. La Lojban. Accessed July 2026. https://mw.lojban.org/papri/Lojban_timeline.
Malotki, Ekkehart. 1983. Hopi Time: A Linguistic Analysis of the Temporal Concepts in the Hopi Language. Berlin: Mouton. https://doi.org/10.1515/9783110822816.
Martinovic, Jasna, Galina V. Paramei, and W. Joseph MacInnes. 2020. Russian blues reveal the limits of language influencing colour discrimination. Cognition 201: 104281. https://doi.org/10.1016/j.cognition.2020.104281.
McElgunn, Hannah. 2024. Benjamin Lee Whorf and Ernest Naquayouma's working relationship: a perspective on linguistic fieldwork in the 1930s. Journal of Anthropological Research 80(4): 383-401. https://doi.org/10.1086/732457.
McWhorter, John H. 2014. The Language Hoax: Why the World Looks the Same in Any Language. Oxford: Oxford University Press.
Mickan, Anne, Maren Schiefke, and Anatol Stefanowitsch. 2014. Key is a llave is a Schlüssel: a failure to replicate an experiment from Boroditsky et al. 2003. Yearbook of the German Cognitive Linguistics Association 2: 39-50. https://doi.org/10.1515/gcla-2014-0004.
Papafragou, Anna, Peggy Li, Youngon Choi, and Chung-hye Han. 2007. Evidentiality in language and cognition. Cognition 103: 253-299. https://doi.org/10.1016/j.cognition.2006.04.001.
Pinker, Steven. 1994. The Language Instinct. New York: William Morrow.
Pullum, Geoffrey K. 1991. The Great Eskimo Vocabulary Hoax and Other Irreverent Essays on the Study of Language. Chicago: University of Chicago Press.
Regier, Terry, and Paul Kay. 2009. Language, thought, and color: Whorf was half right. Trends in Cognitive Sciences 13: 439-446. https://doi.org/10.1016/j.tics.2009.07.001.
Roberts, Seán G., James Winters, and Keith Chen. 2015. Future tense and economic decisions: controlling for cultural evolution. PLoS ONE 10(7): e0132145. https://doi.org/10.1371/journal.pone.0132145.
Robertson, Cole, Seán G. Roberts, Asifa Majid, and Robin I. M. Dunbar. 2025. Language and economic behaviour: future tense use causes less not more temporal discounting. PLoS ONE 20(5): e0317422. https://doi.org/10.1371/journal.pone.0317422.
Samuel, Steven, Geoff Cole, and Madeline J. Eacott. 2019. Grammatical gender and linguistic relativity: a systematic review. Psychonomic Bulletin & Review 26: 1767-1786. https://doi.org/10.3758/s13423-019-01652-3.
Slobin, Dan I. 1996. From "thought and language" to "thinking for speaking." In Rethinking Linguistic Relativity, edited by John J. Gumperz and Stephen C. Levinson, 70-96. Cambridge: Cambridge University Press.
Toki Pona census. 2024. Results of the 2024 Toki Pona census. Updated February 27, 2025. https://tokiponacensus.github.io/results2024/.
Tosun, Sümeyra, Jyotsna Vaid, and Lisa Geraci. 2013. Does obligatory linguistic marking of source of evidence affect source memory? A Turkish/English investigation. Journal of Memory and Language 69: 121-134. https://doi.org/10.1016/j.jml.2013.03.004.
Tse, Chi-Shing, and Jeanette Altarriba. 2008. Evidence against linguistic relativity in Chinese and English: a case study of spatial and temporal metaphors. Journal of Cognition and Culture 8: 335-357. https://doi.org/10.1163/156853708X358218.
Ünal, Ercenur, Adrienne Pinto, Ann Bunger, and Anna Papafragou. 2016. Monitoring sources of event memories: a cross-linguistic investigation. Journal of Memory and Language 87: 157-176. https://doi.org/10.1016/j.jml.2015.10.009.
Ünal, Ercenur, and Anna Papafragou. 2018. Evidentials, information sources, and cognition. In The Oxford Handbook of Evidentiality, edited by Alexandra Y. Aikhenvald, 175-184. Oxford: Oxford University Press. https://doi.org/10.1093/oxfordhb/9780198759515.013.8.
Winawer, Jonathan, Nathan Witthoft, Michael C. Frank, Lisa Wu, Alex R. Wade, and Lera Boroditsky. 2007. Russian blues reveal effects of language on color discrimination. Proceedings of the National Academy of Sciences 104: 7780-7785. https://doi.org/10.1073/pnas.0701644104.
Witzel, Christoph, and Karl R. Gegenfurtner. 2011. Is there a lateralized category effect for color? Journal of Vision 11(12): 16. https://doi.org/10.1167/11.12.16.
Wolff, Phillip, and Kevin J. Holmes. 2011. Linguistic relativity. WIREs Cognitive Science 2: 253-265. https://doi.org/10.1002/wcs.104.
Phi sources
The Phi material in this chapter comes from the number and evidential rulings in canon.md, the number chapters under manual/part4_grammar/ch12_numbers/, and the evidential chapters under manual/part4_grammar/ch16_evidentiality/. Its boundary on claims follows project/development_protocol.md, especially the rule that design intentions must be tested in use rather than presented as guarantees.