T2 · Preprint version 2 · Derivative text-only reading copy. Cite https://doi.org/10.31235/osf.io/e9qw5_v2. Supplied figure legends are retained; omitted figures are not reconstructed. This page is generated from the consolidated manuscript; guides do not replace it.

Sections

Paper T2 — Philosophy as Cognitive Assay: Measuring the Delegation Legitimacy Boundary in AI-Assisted Knowledge Work

Preprint version: 2
Canonical DOI: https://doi.org/10.31235/osf.io/e9qw5_v2

Kengo Tomita Institute of Technology, Shimizu Corporation


Abstract

This article operationalizes the concept of a delegation legitimacy boundary — the structural line along which human judgment can and cannot be legitimately delegated to artificial intelligence — and proposes a minimal scoring protocol for locating it within any knowledge-work task. Building on the Mekiki framework (Tomita 2026), which distinguished specification — the task-defining substrate of domain expertise — from externalization cost — the technical barrier that AI selectively removes — the present article derives a further decomposition within specification itself. Drawing on case-study evidence in which recorded specification judgments contain separable factual components (Sein-type: what the data structure affords, how users will perceive a display) and value-laden components (Sollen-type: what information ought to be excluded, which priority ranking is appropriate), the article grounds this distinction in Hume's is–ought separation and Kant's Sein/Sollen architecture, and redeploys it as the axis of a cognitive assay — a measurement system in which philosophical categories serve not as normative prescriptions but as diagnostic coordinates. The resulting five-step scoring protocol assigns Sein-type components to AI evaluation and Sollen-type components to domain-expert evaluation; this asymmetry is not a design choice but a structural consequence of the boundary itself, which holds as long as legitimacy over value judgments remains institutionally human-attributed. Most individual judgments are hybrid, carrying both components in varying ratios; the protocol therefore yields ratio profiles rather than binary classifications. As a first application, the article re-describes the four processes of Nonaka and Takeuchi's Socialization–Externalization–Combination–Internalization (SECI) model through the assay, deriving as an analytic consequence the finding that AI acceleration of Externalization and Combination shifts the effective rate-limiting stage to Socialization and Internalization — both human-limited cognitive and social processes that cannot be accelerated by AI investment alone.

Keywords: human judgment, delegation legitimacy, cognitive assay, Sein/Sollen, tacit knowledge, AI-assisted knowledge work


1. Introduction

Something that could not previously be measured is becoming measurable. Large language models have compressed the cost of converting expert knowledge into formal output — code, reports, draft designs — to a degree that would have seemed implausible a decade ago. Yet this compression has not rendered expertise redundant. On the contrary, it has exposed the judgments that expertise carries: what should be built, what should be excluded, and by what criteria quality should be assessed. Practitioners across domains now confront, often for the first time, the question of where their own contribution begins and the machine's contribution ends. Stader (2024), in adjacent philosophical work, captured the asymmetry precisely: judgment is not replaced by calculation; judgment is what gives calculation its meaning and purpose. The Mekiki framework (Tomita 2026) gave this asymmetry an operational structure by decomposing the Externalization process into two conditions — specification (the domain expertise that determines what should be built) and externalization cost (the technical barriers that AI selectively removes) — and demonstrating through case-study and domain-ablation evidence that the two conditions vary independently. What the framework left unexamined was the internal structure of specification itself.

The question of what lies inside that specification has, meanwhile, been approached from the direction of legitimacy. A growing body of work has shifted the normative criteria for AI-assisted decisions from explainability — can the system's reasoning be rendered transparent? — toward justifiability and contestability: can the decision be warranted on grounds that those affected would recognize as legitimate (Henin and Le Métayer 2022; Beckman et al. 2024)? The distinction matters because explanation is primarily descriptive — it answers "how did the system arrive at this output?" — while justification is inescapably normative: it must answer "on what authority, and toward what end?" Responsibility-gap analyses have shown that when this normative dimension is elided, accountability diffuses into organizational structures where no identifiable agent bears the evaluative burden (Dastani and Yazdanpanah 2023; Taylor 2025). Crucially, this traceability is not only a matter of accountability — which can diffuse across institutional structures — but more specifically of answerability: the non-transferable relation in which the warrant of a judgment remains anchored in a particular agent. What these analyses share is the recognition that some component of decision-making resists delegation not because AI lacks the computational capacity to produce an output, but because the output's warrant must be traceable to a source whose legitimacy is socially recognized.

A parallel conversation has been unfolding around the structural fate of tacit knowledge. Gill (2023), introducing a special issue of AI & Society, placed tacit knowledge — the dimension of knowing that resists full articulation (Polanyi 1966) — at the center of any serious account of AI futures, arguing that what cannot be articulated is precisely what is most at risk of being overlooked in systems designed around the articulable. Collins (2025) sharpened the point: large language models possess no moral compass — they cannot distinguish between what is factual and what is fabricated, because the distinction requires evaluative commitments that lie outside statistical pattern-matching. Ferdman (2025) extended the diagnosis from individual cognition to structural conditions, arguing that AI-induced deskilling is not a failure of personal discipline but a predictable consequence of capacity-hostile environments — institutional arrangements that systematically degrade the conditions under which expertise is acquired and exercised. These contributions converge on a shared concern: that AI systems, in accelerating the articulable dimensions of knowledge work, may simultaneously be eroding the conditions that sustain the inarticulate dimensions on which quality ultimately depends.

Yet for all the sophistication of these diagnoses, a specific intermediate layer remains absent. Existing scholarship has established that human oversight is necessary, that democratic legitimacy matters, that responsibility must be assignable, and that audit mechanisms require development (Manheim et al. 2025; Concannon and Tomalin 2024). What is missing is not another argument for the importance of human judgment — that case has been made convincingly and repeatedly — but a measurement system: a structured method for decomposing any given judgment into components, determining which components carry what kind of evaluative weight, and assigning each to an appropriate scorer. The gap, in other words, is not empirical but conceptual-operational. The problem has been diagnosed; the instrument for acting on the diagnosis has not yet been proposed in consolidated form. The present article does not claim untouched territory; it consolidates a missing intermediate layer between existing diagnoses and partial implementations. A recent Japanese-language textbook integrating AI technology, philosophy of mind, and social governance offers an integrative account of these domains (Taniguchi et al. 2026); the present article asks the complementary question of how to operationally measure the delegation legitimacy boundary.

The present article proposes such an instrument. Taking the specification identified by the Mekiki framework as its analytical substrate, it derives a decomposition into Sein-type (factual) and Sollen-type (value-laden) components, grounding the distinction in Hume's is–ought separation (1739) and Kant's critical architecture (1781, 1785). It then redeploys these philosophical categories not as normative prescriptions — not as claims about what AI should or should not do — but as the axes of a cognitive assay: a diagnostic system in which the categories serve as measurement coordinates. The term "assay" is used here in its analytical sense: a structured test that makes a composite judgment legible by separating its components and identifying where AI can reduce externalization cost without crossing the boundary of legitimate delegation. The five-step scoring protocol that results assigns Sein-type components to AI evaluation and Sollen-type components to domain-expert evaluation. This asymmetry is not imposed by design but follows structurally from what the article terms the delegation legitimacy boundary — the line beyond which AI may generate outputs but cannot serve as the source of their legitimacy. The boundary is not a capability claim; it is a legitimacy claim, and it holds as long as the authority to make value judgments remains institutionally attributed to human agents.

As a first application, the article re-describes the four processes of Nonaka and Takeuchi's (1995) SECI model — Socialization, Externalization, Combination, and Internalization — through the lens of the assay. This re-description is not an exegesis of SECI but a constructive refinement: it asks what becomes visible when each process is decomposed into Sein-type, Sollen-type, and externalization-cost components. The analytic consequence is that AI's acceleration of E (Externalization) and C (Combination) — the two processes most saturated with externalization cost — shifts the effective rate-limiting stage to S (Socialization, where Sollen-type specification is primarily updated through shared experience) and I (Internalization, where Sein-type specification is primarily absorbed into individual practice). Both remain human-limited processes. The implication for the tacit-knowledge concerns raised by Gill, Collins, and Ferdman is direct: the processes that cannot be accelerated by AI investment alone are those that sustain the evaluative capacities on which knowledge quality depends. The article thus bridges the legitimacy discourse, the tacit-knowledge discourse, and the measurement discourse by offering a single structural vocabulary in which all three become aspects of the same underlying decomposition.

2. The internal structure of specification: deriving the Sein/Sollen decomposition

2.1 Three concepts, not two

The Mekiki framework (Tomita 2026) distinguished three concepts that are routinely conflated in discussions of AI-assisted knowledge work. Domain expertise is a resource that a practitioner possesses — accumulated through training, practice, and exposure to a field's materials and norms. Specification cost is a property of the task: the degree to which the task demands that resource. And specification — the term used on Figure 1 of the original article as "substrate" — is what results when domain expertise is invested in a particular task: the concrete judgments about what should be built, what should be excluded, and by what criteria quality should be assessed. The present article takes the first two concepts as given. Its question is narrower: what is the internal structure of specification itself? When a domain expert invests expertise in a task, do the resulting judgments all share the same epistemic character, or do they contain components of fundamentally different kinds?

2.2 Derivation from the Mekiki case study

The Mekiki article recorded, in granular detail, the specification judgments that a domain expert made during the development of an EV charging application. Revisiting those records, two distinguishable kinds of judgment emerge.

One kind is factual. The domain expert judged that a current-based data structure reflects charging state more accurately than a kilowatt-hour display. That the identifier architecture of a particular network operator follows a specific hierarchical pattern. That a given display condition would produce user misrecognition. These judgments concern what is the case — how a data structure behaves, what a user will perceive, what a system affords. Given sufficient data and domain-relevant training, an AI system could in principle arrive at the same determinations.

The other kind is value-laden. The domain expert judged that real-time charger availability data — whether a charging station is currently in use — should not be displayed, even if technically obtainable: a non-adoption decision grounded in the judgment that unreliable availability information would undermine the application's purpose as a pre-trip planning tool. That a particular filter condition was appropriate for the intended user. That one information hierarchy was preferable to another. These judgments concern what ought to be the case — what serves the user well, what respects the design's normative intent, what information to withhold. An AI system can generate such outputs, but it cannot supply the warrant: the evaluative commitment that makes the judgment legitimate must originate from an agent whose authority to make it is socially recognized.

This article treats this two-way distinction as a minimal decomposition. The claim is not that every judgment falls cleanly into one category or the other. Most judgments in practice are hybrid: they carry both factual and value-laden components in varying proportions. Williams (1985) called such composites "thick concepts" — terms like cruel, courageous, or negligent that simultaneously describe and evaluate. The minimal decomposition does not deny hybridity; it asserts that within any hybrid judgment, it is analytically productive to ask how much of its weight rests on factual grounding and how much on evaluative commitment. The scoring protocol introduced in Section 3 operationalizes this question as a ratio profile rather than a binary classification.

2.3 Philosophical foundations

The distinction just derived from case-study evidence is not new. It is, in fact, among the oldest in Western philosophy. Hume (1739) observed that no amount of descriptive premises about what is the case can, by themselves, yield a conclusion about what ought to be the case — the is–ought gap. Kant systematized the distinction as Sein (being, the domain of theoretical reason) and Sollen (ought, the domain of practical reason), arguing that the two are governed by different rational faculties and irreducible to each other (Kant 1781, 1785). The present article adopts the Kantian terminology because it names the distinction with greater precision than the English is–ought pairing, and because the Neo-Kantian tradition — particularly Weber's insistence on value-freedom in social science (Weber 1904) — makes the distinction legible within the social-theoretical vocabulary familiar to this journal's readership.

Two subsequent contributions deepen the philosophical foundation. As Williams (1985) showed for thick concepts, clean decomposition into pure Sein and pure Sollen is not always possible — a constraint the present protocol addresses through ratio profiles rather than binary classifications. Nagel (1979, Ch. 9) went further, arguing that value is irreducibly fragmented — that distinct types of practical reason (obligations, rights, utility, perfectionist ends, personal commitments) cannot be reduced to a single scale or to agent-neutral terms. His conclusion — that good judgment without complete justification is the best available response to this fragmentation — is a philosophical anticipation of what the present framework operationalizes as specification: the practitioner's capacity to make situated evaluative judgments that resist algorithmic reduction. Nagel (1986, Ch. VIII–IX) extended the argument: agent-relative reasons — reasons that depend on the particular perspective, commitments, and situation of the agent — are irreducible to agent-neutral reasons accessible from any viewpoint. This irreducibility provides the deepest philosophical justification for the scoring asymmetry introduced in Section 3: Sollen-type components require evaluation by an agent whose perspective is constitutively relevant to the judgment, not by a system that operates from a view from nowhere. Nagel observed that what is needed is a method for breaking up practical problems so as to see which considerations are relevant and how they bear on a conclusion, but did not himself construct such a method. The present protocol is, in part, a response to that forty-year-old request.

The delegation legitimacy boundary, introduced in Section 1, runs along this distinction. Sein-type specification is, in principle, gradually absorbable into AI processing: as training data and domain models improve, factual determinations that once required human expertise become amenable to automated evaluation. The boundary is permeable on the Sein side — what counts as "requiring human judgment" shifts as AI capabilities advance. This permeability concerns the automation of factual evaluation, not the transfer of normative authority. Sollen-type specification, by contrast, resists this absorption not because AI cannot produce value-laden outputs — it manifestly can — but because the legitimacy of such outputs cannot be grounded in the system itself. The claim is not about capability but about warrant. An AI system that recommends withholding certain data from users may be producing exactly the judgment a domain expert would make; but the authority to make that judgment — the right to impose a normative commitment on others — must be traceable to a human source whose legitimacy is institutionally recognized.

Combining this with the Mekiki framework's two-condition structure yields a three-layer description of knowledge work under AI assistance (Fig. 1):

  1. Externalization cost — technical barriers to converting expertise into formal output. AI removes or substantially compresses this layer.
  2. Sein-type specification — factual judgments within the invested expertise. Gradually learnable by AI; over time, some determinations may become routinized and handled as part of externalization-cost processing.
  3. Sollen-type specification — value-laden judgments within the invested expertise. The delegation legitimacy boundary sits here: AI may produce these outputs, but cannot serve as the source of their legitimacy.

A clarification is necessary here. Sollen-type specification — the practitioner's capacity to make value judgments — should not be confused with Sollen as an environmental normative pressure (e.g., institutional norms, regulatory requirements, or social expectations that constrain action). The distinction mirrors the one the Mekiki framework drew between specification (a resource the practitioner possesses) and specification cost (a demand the task imposes). Confusing the two would invite the misreading that Sollen-type specification is an external burden to be removed, when it is in fact the evaluative capacity that gives knowledge work its direction.

2.4 Generalization beyond Externalization

One further step is required before the decomposition can serve as the axis of an assay. The Mekiki framework derived the specification/externalization-cost distinction within the Externalization process of Nonaka and Takeuchi's SECI model. But nothing in the distinction itself is specific to Externalization. Böhm and Durst (2026), in their GRAI framework, added a human/machine agent dimension to SECI's four modes; the Mekiki article noted that the externalization-cost/specification distinction applies orthogonally within any SECI mode and any GRAI interaction field. The present article extends that observation: not only the two-condition structure but also the Sein/Sollen decomposition within specification cuts across all four processes. Every SECI process — Socialization, Externalization, Combination, Internalization — involves activities that carry both factual and value-laden components, in different proportions. Socialization (shared experience) is saturated with Sollen-type judgments about what matters, what to attend to, what norms to absorb. Internalization (learning by doing) is weighted toward Sein-type absorption of established knowledge into individual practice, though it also involves normative recalibration. The Sein/Sollen distinction is orthogonal to the SECI process distinction: it applies within any process, not between them.

This orthogonality is what makes the distinction usable as an assay axis. If Sein and Sollen were simply another way of labeling E and S, the decomposition would add no analytical power. Because they cut across SECI processes at right angles, they enable a matrix analysis — each process decomposed by judgment type — that reveals structural asymmetries invisible from within either framework alone. Section 4 carries out this analysis.

The level of analysis throughout this article is the structure of knowledge activities — the process level. The framework does not make claims about individual cognition (how a single expert internally represents Sein and Sollen), nor about organizational design (how institutions should be restructured), nor about regulatory policy (what laws should govern AI delegation). Extensions to those levels are possible and, the author believes, productive — but they belong to the domain experts of those respective fields. This article is a constructive refinement of SECI, not an exegesis: it asks what becomes visible when SECI's processes are examined through the Sein/Sollen lens, and it limits its claims to what that examination reveals.

3. Philosophy as cognitive assay: operationalization

3.1 Why the assay conditions are now met

In pharmacological assay design, a compound's activity can only be measured when background noise is sufficiently attenuated. If the noise floor is too high, the signal of interest — however strong — remains undetectable. The same structural logic applies here. Under conventional conditions of knowledge work, specification and externalization cost were entangled and inseparable (Tomita 2026): the effort of articulating an idea was indistinguishable from the expertise that gave the idea its content. AI's compression of externalization cost has substantially lowered this noise floor. Specification has not changed in character, but it has become isolable as a variable — foregrounded against a backdrop of dramatically reduced conversion friction. The Mekiki article's case study and domain-ablation experiment captured this transition in progress: the same specification, routed through different AI systems, produced outputs whose quality varied as a direct function of the specification invested, with externalization cost held approximately constant across conditions. Kusumegi et al. (2025), analyzing two million academic manuscripts, observed the complementary pattern at scale: AI-assisted writing improved the linguistic quality of submissions (externalization cost reduced) while acceptance rates remained unchanged (specification quality unaffected) — the two components varying independently in a population-level natural experiment. These conditions — specification foregrounded, externalization cost attenuated, independent variation empirically documented — are the conditions under which an assay becomes viable.

3.2 The scoring protocol

The unit of analysis is the individual feature decision: a single judgment recorded in the design process — a data structure choice, a display condition, a filter rule, a non-adoption decision. The Mekiki case study contains approximately thirty such decisions; Section 2.2 showed that they contain separable Sein-type and Sollen-type components.

The protocol proceeds in five steps.

Step 1. Feature-decision extraction (AI). The AI system identifies discrete judgments from the design record — requirements documents, dialogue logs, commit histories, or equivalent artifacts. This is a Sein-type operation: pattern recognition over documented material.

Step 1b. Gap detection (domain expert). The expert reviews the extracted list and adds judgments that the AI missed — in particular, non-adoption decisions (features deliberately excluded) and implicit prioritizations (orderings that were never explicitly stated). The Mekiki case study's decision to exclude real-time charger availability data is an example: the AI could not extract it because the feature was never built, and its absence from the record was itself the judgment. Step 1b corresponds to what the present article operationalizes as a high-order layer of specification cost — detecting the absence of something that ought to be present. AI can extract existing decisions; it cannot detect the non-existence of what should exist.

Step 2. Operation-type classification (AI). Each feature decision is labeled by operation type — detect, prioritize, exclude, or justify — as a descriptive tag independent of the scoring that follows. A single decision may carry multiple labels (e.g., detect/exclude). All four types can appear in both the Sein and Sollen rows: an exclusion based on data quality is Sein-type; an exclusion based on user values is Sollen-type.

Step 3. Sein-row scoring (AI). For each feature decision, the AI evaluates the factual component on a three-point scale: 0 (not detected / inaccurate), 1 (partial / ambiguous), 2 (accurate / appropriate). The anchor for this row is data-grounded: the AI assesses whether the factual determination is consistent with available evidence.

Step 4. Sollen-row scoring (domain expert). For each feature decision, the domain expert evaluates the value-laden component on the same scale. The anchor here is normative: does the judgment serve the design's purpose, protect the user's interests, reflect appropriate professional standards? The order in which the expert scores individual decisions is left to the expert's discretion. Because Sollen-type judgments lack natural ground truth, reliability is established through independent scoring by multiple experts followed by adjudication of disagreements.

Step 5. Profile analysis (AI aggregation + expert interpretation). Scores are not summed into a single index. Sein-type quality and Sollen-type quality are incommensurable dimensions; their addition would be meaningless — adding factual adequacy to normative warrant would produce a pseudo-precision that conceals rather than clarifies the task structure — and this incommensurability is itself a structural consequence of the two-component decomposition. Instead, the protocol reports ratio profiles: the Sein-row mean and variance, the Sollen-row mean and variance, and the distribution of hybrid judgments across the task. Aggregation by operation type (e.g., average Sein/Sollen ratio for exclude-type decisions versus detect-type decisions) provides a secondary layer useful for cross-task comparison. The expert's role at this step is interpretive: what does the profile reveal about where human judgment is most critical in this task? At the lowest level, each feature decision is itself a local Sein/Sollen profile — pure Sein, pure Sollen, or hybrid — and the task-level profile is the distribution of these local profiles.

Two structural features of the protocol deserve emphasis. First, the scorer asymmetry — AI for Sein rows, domain expert for Sollen rows — is not a design preference but a consequence of the decomposition derived in Section 2. Sein-type judgments admit data-grounded verification; Sollen-type judgments require a legitimacy source that AI cannot supply. Second, the workflow itself embodies the asymmetry it measures: Steps 1 through 3 are fully automatable (externalization-cost processing); the expert's decisive scoring interventions occur at Steps 1b and 4 (specification in its Sollen dimension). A protocol for measuring the delegation legitimacy boundary that was itself fully delegable would be self-defeating.

The asymmetry also has an institutional function. Sein-row scoring by AI is not neutral in an absolute sense — model architecture, training data, and reward design carry their own biases — but it can provide a procedural check on interpersonal distortions that often affect human evaluation, including status protection, face-saving, motivated reasoning, and local political pressure. The point is not to treat AI as an authority, but to use it as a provisional factual scorer whose outputs remain auditable and contestable. This procedural function is one structural reason why the asymmetry is institutional rather than merely technical.

Expert selection. Domain-specific criteria are a matter for each field, but scorers should hold accountability for the relevant judgment type, occupy an evaluative role, and, where possible, score independently with adjudication.

Validity conditions and failure modes. The assay's predictive validity rests on evidence that specification quality predicts output quality independently of externalization-cost quality — a condition supported by the Mekiki case study and by Thorgeirsson et al. (2026), who found that computer-science achievement predicted AI-assisted programming performance ("vibe coding") (r = 0.39) while LLM usage frequency correlated negatively (r = -0.26). Discriminant validity requires that Sein-type and Sollen-type scores vary independently; the domain-ablation experiment (Section 3.3 below) and Kusumegi et al.'s population-level data both confirm this independence. The assay admits two principal failure modes. False positives occur when high externalization-cost quality masks the absence of specification — fluent prose concealing hollow content, or sycophantic AI output (Chandra et al. 2026) reinforcing a user's Sollen commitment without independent warrant. False negatives occur when residual externalization-cost friction renders genuine specification invisible — an expert whose judgment is sound but whose articulation is poor receives artificially low scores.

3.3 Worked example: a minimal demonstration

To illustrate the protocol's separating power, two judgments from the Mekiki case study are scored here, selected as analytically clean representatives — one Sein-dominant, one Sollen-dominant — from the case study's larger inventory of feature decisions; the full inventory is available in the companion article (Tomita 2026).

Judgment A (Sein-dominant): "A current-based data structure reflects charging state more accurately than a kilowatt-hour display." Sein row: 2 (verifiable by data comparison). Sollen row: 0 (no value commitment involved). Operation type: detect.

Judgment B (Sollen-dominant): "Real-time charger availability data should not be displayed, even if technically obtainable, because unreliable availability information would undermine the application's purpose as a planning tool." Sein row: 1 (the data could in principle be retrieved). Sollen row: 2 (a value judgment about what information to withhold from users). Operation type: exclude.

Under the domain-ablation conditions reported in the companion article — where AI systems built the same application without domain expertise — the Sein row retained partial function (data retrieval succeeded) but the Sollen row dropped systematically toward zero: no exclusion judgments, no user-relevance filtering, no prioritization, and, in some cases, even factual semantics were misread. The two components separate on the rubric, and the scorer asymmetry follows as a structural necessity. Section 4 applies the assay at a larger scale, re-describing the full SECI model through the Sein/Sollen decomposition.

Criterion contamination. This worked example is drawn from the theory-builder's own case study. Independent validation through blind scoring by domain experts who are naive to the framework would constitute a necessary external-validity test; the present demonstration establishes only that the protocol can be applied and that it separates the two components as predicted.

3.4 Philosophical frameworks as candidate assay axes

The Sein/Sollen distinction is the minimal axis required for the present assay, but it is not the only philosophical distinction that could serve as a measurement coordinate. Aristotle's phronesis — practical wisdom, the capacity to judge rightly in particular circumstances (Aristotle, Nicomachean Ethics VI) — provides a candidate axis for evaluating context-dependent Sollen-type specification: not whether a judgment conforms to a rule, but whether it reflects appropriate discernment of the situation. Kant's distinction between determinative judgment (applying a given universal to a particular case) and reflective judgment (finding a universal for a given particular) maps onto the AI-amenable/AI-resistant divide from a different angle: determinative judgment is structurally closer to Sein-type evaluation; reflective judgment to Sollen-type. Wright et al. (2026), who used large language models as zero-shot personality assays, demonstrated that AI can function as a measurement instrument for psychological constructs — a precedent for the broader claim that AI-mediated evaluation can serve as a cognitive assay, provided the construct being measured is carefully specified.

A further candidate axis lies within Sollen-type specification itself. Nagel's (1986) taxonomy of agent-relative and agent-neutral reasons (cf. Parfit 1984) suggests that Sollen-type components may admit sub-classification — utilitarian considerations (agent-neutral), autonomy-based commitments (agent-relative), deontological obligations (agent-relative but universalizable). The present protocol treats Sollen-type specification as an undifferentiated category by design: this is the minimal decomposition required for the assay to function, and sub-classification is deferred. What matters for the present argument is Nagel's demonstration that none of these sub-categories is reducible to agent-neutral terms — and therefore none is reducible to Sein-type evaluation. The irreducibility holds regardless of how finely the Sollen side is carved.

The same distinction also has a dynamic implication: when evaluative commitments associated with Nagel's reasons become embedded in social practice, their institutionalization should be described as a state change within Sollen-type specification, not as a conversion of Sollen-type specification into Sein-type specification. The distinction introduced here is therefore dynamic rather than taxonomic. At the level relevant to delegation, agent-relative evaluative commitments sustained through practice and oriented toward ideals of excellence constitute a Do-type state of Sollen-type specification (from Japanese , practice-sustained engagement; elaborated in Section 6.5). Through social processes of legitimacy acquisition, such commitments can become institutionalized and textualized — codified in regulations and professional standards, or otherwise represented in training data — thereby taking Sin-type form, where Sin abbreviates Sollen institutionalized: a state in which normativity operates primarily through compliance and violation avoidance. In this form, Sin-type material is agent-neutralized and more readily absorbed by AI as reproducible textual regularity. This is the limited, operational sense in which Sin-type material behaves structurally like Sein-type specification: AI can reproduce the codified settlement, but it does not inherit the legitimacy by which that settlement became authoritative. Conversely, when institutional conditions shift, when codified rules become contested, or when their application becomes underspecified, Sin-type material reopens into Do-type judgment. Reasons of autonomy therefore remain the clearest limiting case for legitimate delegation: their outputs may later be institutionalized, but the evaluative standard originates within the agent and cannot itself be absorbed as reproducible textual regularity. This dynamic distinction is introduced as an interpretive extension, not as an additional scoring dimension in the present protocol. These axes are not developed further here; they are noted as coordinates available for future assay designs.

4. First application: re-describing SECI through the assay

The cognitive assay developed in Sections 2 and 3 was derived from the Externalization process. But as Section 2.4 established, the Sein/Sollen distinction is orthogonal to the SECI process distinction: it applies within any process, not between them. This section carries out the matrix analysis that orthogonality permits, decomposing each of Nonaka and Takeuchi's (1995) four knowledge-creation processes into its externalization-cost, Sein-type specification, and Sollen-type specification components. The re-description is not an exegesis of the original model but a constructive refinement: it asks what structural asymmetries become visible when SECI is examined through the assay's lens.

4.1 Externalization: the Mekiki framework's direct domain

Externalization — converting tacit knowledge into explicit form — is the process where AI's impact is most immediately visible and where the Mekiki framework was originally derived. The Sein/Sollen decomposition reveals why. Externalization is dominated by externalization cost: the technical effort of translating what one knows into documents, code, diagrams, or formal specifications. AI compresses this cost dramatically. Sein-type specification participates as the factual content being converted — the data structures, causal relationships, and empirical regularities that enter the formal output. Sollen-type specification, however, remains invariant through the Externalization process: the judgment of what ought to be externalized, which priority to assign, which information to withhold, does not change merely because the conversion has become easier. Kusumegi et al. (2025) documented this invariance at population scale: AI improved the externalization-cost dimension of two million manuscripts while leaving acceptance rates — a proxy for specification quality — unchanged.

4.2 Combination: AI's strongest domain

Combination — integrating, sorting, and recategorizing explicit knowledge — is the process most saturated with externalization cost and correspondingly most amenable to AI acceleration. Database integration, cross-referencing of literature, pattern detection across codified sources: these are operations on already-formalized knowledge, and AI performs them at a speed and scale that no human team can match. Sein-type specification is low because the knowledge being combined has already been formalized; its factual adequacy was established during prior Externalization. Sollen-type specification enters only at evaluation: which combinations are meaningful, which juxtapositions merit attention, which syntheses serve a purpose. The Combination process itself is largely automatable; the judgment of what its outputs mean is not.

4.3 Socialization: the domain of Sollen-type specification

Socialization — the sharing of tacit knowledge through shared experience, co-practice, and what Nonaka and Takeuchi called "intellectual combat" — is the process where Sollen-type specification is most concentrated. This is where value judgments are formed, contested, and revised: what matters in this domain, what standards of quality to maintain, what norms to absorb or challenge. The sharing of factual knowledge (Sein-type) can increasingly be supported by documentation or AI-mediated translation, but the non-trivial revision of value commitments — the moment when a practitioner's evaluative framework shifts through encounter with a different perspective — requires human deliberation. AI can participate in Socialization, generating alternative perspectives or surfacing overlooked considerations; but it cannot serve as the source of legitimacy for the resulting normative commitments. The delegation legitimacy boundary is most visible here. Scharfmann, Marx, and Fleming (2025), analyzing 682,000 researchers who moved between academia and industry, found that cross-domain co-location — a form of institutionalized Socialization — systematically increased the novelty and impact of subsequent work. The pattern they documented is consistent with the mechanism the assay would lead one to expect: exposure to different Sollen-type commitments (different standards, different priorities) catalyzes specification refinement in ways that no amount of within-domain Combination can achieve.1

4.4 Internalization: AI as catalyst, not substitute

Internalization — absorbing explicit knowledge into individual tacit practice — is weighted toward Sein-type specification: the practitioner learns established facts, procedures, and patterns until they become second nature. AI can serve as a catalyst in this process, providing adaptive instruction, worked examples, and immediate feedback that accelerate the absorption of codified knowledge. But Internalization also involves normative recalibration: the practitioner does not merely acquire facts but develops a sense of what matters within a domain — an evaluative orientation that emerges through practice rather than instruction. This Sollen-type dimension of Internalization is what distinguishes expertise from mere knowledge: the expert not only knows the facts but has internalized the evaluative commitments that govern their application. AI can support the Sein-type dimension of Internalization; it cannot substitute for the experiential formation of evaluative judgment.

4.5 The rate-limiting shift: an analytic consequence

Table 1 summarizes the decomposition; Fig. 2 provides a visual representation of the rate-limiting shift.

Table 1. SECI processes decomposed by the cognitive assay's three components.

SECI process Ext. cost Sein-type Spec. Sollen-type Spec. AI effect
S (Socialization) Low Medium High Low (no legitimacy source)
E (Externalization) High Medium Low (invariant) Maximal (cost compression)
C (Combination) Highest Low Low (evaluation only) Maximal
I (Internalization) Medium High Medium (emerging Sollen) Medium (catalytic support)

The pattern is structural. AI's maximal effect falls on the two processes — E and C — where externalization cost dominates. Its minimal effect falls on S, where Sollen-type specification dominates and the delegation legitimacy boundary is most active. I occupies an intermediate position: AI can catalyze the Sein-type absorption but cannot substitute for the experiential formation of evaluative commitments.

The consequence for the SECI spiral's dynamics follows analytically. When two processes in a sequential cycle are dramatically accelerated while the other two remain at human speed, the effective rate-limiting stage shifts to the unaccelerated processes. In reaction-pathway kinetics, the slowest step in a sequential pathway determines the overall throughput. Applied to SECI: AI acceleration of E and C shifts the effective bottleneck to S and I. This is not an empirical finding but an analytic consequence of the matrix above — a structural prediction that follows from the assay's decomposition.

The matrix values in Table 1 are theoretical assignments — ordinal characterizations derived from the assay's logic, not measured quantities. Empirical calibration of these values across domains is a matter for future research. What the matrix provides is not measurement but structural visibility: it makes explicit which cells carry the specification that AI cannot supply and which carry the cost that AI compresses. The rate-limiting prediction follows from the structure regardless of the precise values in each cell.

One further structural feature deserves note. Sein-type specification undergoes a spiral transition across SECI cycles: what begins as novel factual insight (high specification cost) is, through repeated Externalization and Combination, gradually codified and absorbed into routine processing — eventually becoming part of the externalization-cost infrastructure that subsequent cycles take for granted. Sollen-type specification does not undergo this transition. Value commitments may be revised through Socialization, but they do not migrate into automated processing; they remain institutionally tied to human evaluative authority. This asymmetry reinforces the rate-limiting prediction: AI's domain expands along the Sein axis but encounters a structural boundary along the Sollen axis.

5. Convergent evidence

The assay proposed in this article is a conceptual-operational contribution, not an empirical study. However, the decomposition it introduces generates structural predictions, and several independent bodies of evidence — none designed to test the framework — converge on patterns consistent with those predictions. This section organizes that evidence by validity type.

5.1 Predictive validity

The assay predicts that specification quality determines output quality independently of externalization-cost quality. Three independent observations are consistent with this prediction at different scales. The Mekiki case study (Tomita 2026) documented that specification-present conditions produced a deployable application within ten days, while specification-absent conditions (domain-ablation) produced functionally complete but domain-inappropriate outputs — the same AI, different specification, different quality. Thorgeirsson, Weidmann, and Su (2026), in a controlled study of one hundred computer-science (CS) students, found that prior CS achievement predicted vibe-coding performance — a direct individual-level analog of specification predicting output quality. At industrial scale, Anthropic's Project Glasswing (Anthropic 2026a) reported that Claude Mythos Preview autonomously discovered thousands of zero-day vulnerabilities across all major operating systems and web browsers — a saturation of the Externalization process in cybersecurity. The resulting bottleneck was not further vulnerability discovery but triage, remediation priority, and coordination among partner organizations — the kinds of Socialization- and Internalization-weighted activities that the assay's matrix (Table 1) would lead one to expect as rate-limiting under AI acceleration of E and C.

5.2 Discriminant validity

The assay predicts that Sein-type and Sollen-type components vary independently — that improving one does not automatically improve the other. Kusumegi et al. (2025), analyzing two million academic manuscripts, provided the largest available test of this prediction: AI-assisted writing substantially improved linguistic quality (externalization cost reduced) while journal acceptance rates remained unchanged (specification quality unaffected). The two components varied independently in a population-level natural experiment. Thorgeirsson et al. (2026) confirmed the same independence at the individual level: writing ability (an externalization-cost proxy) ceased to predict vibe-coding performance after controlling for cognitive ability, while CS achievement (a specification proxy) remained a significant predictor — the two constructs are statistically separable. The Mekiki domain-ablation experiment demonstrated the extreme case: AI systems with full externalization-cost processing capability but zero domain specification produced outputs in which the Sollen row dropped to zero while the Sein row retained partial function — a counterfactual demonstration that the two components can be dissociated.

5.3 Convergent and ecological validity

Two further studies provide convergent support from different disciplinary angles. Kubota et al. (2026), in an LLM-assisted replication study, observed a gradient from verifiability to generalizability in the dimensions of scientific work that AI could replicate — a gradient that maps onto the externalization-cost-to-specification continuum, providing methodological support for the assay's operational structure. Scharfmann, Marx, and Fleming (2025), whose cross-domain co-location findings were discussed in Section 4.3, provide ecological support from the opposite direction: the pattern they documented is consistent with the assay's prediction that S-process investment refines specification in ways that C-process activity cannot. Huang (2025), writing independently on content moderation, proposed a parallel shift from accuracy to legitimacy as the evaluative framework for LLM-assisted decisions — a domain-specific instance of the general decomposition proposed here. Klowden and Tao (2026), analyzing AI-generated mathematical output, identified what they call an "unprecedented separation of form and thought" — the observation that AI can produce formally correct derivations that lack the conceptual coherence a mathematician would recognize. What Tao describes informally as a "smell test" for mathematical argument is, in the present framework's terms, a Sollen-type specification judgment: an evaluative assessment that cannot be reduced to formal verification. The structural tension they document — between externalization-cost quality (formal correctness) and specification quality (mathematical insight) — is an independent observation of the same decomposition in a domain far removed from the case study that generated the present framework.

Additional convergent signals come from three further fields; none tests the assay directly, but each registers the same asymmetry in a neighboring vocabulary. Dell'Acqua et al. (2026), in a field experiment with 758 knowledge workers, found that AI assistance improved productivity and quality for tasks within the "jagged technological frontier" but reduced accuracy by 19 percentage points for tasks beyond it — a boundary that maps onto the externalization-cost/specification divide. Tang et al. (2025) showed that once rewards become unverifiable, reinforcement learning for language models requires different objectives and weaker evaluation guarantees — a machine-learning indication that Sein-type evaluation scales more readily than Sollen-type evaluation. Dorner, Nastl, and Hardt (2025) proved mathematically that AI evaluators at the evaluation frontier cannot achieve efficiency gains exceeding a factor-two improvement over the ground-truth baseline — a formal bound on the automation of judgment that is consistent with the scoring asymmetry proposed here. These three studies, drawn from organizational science, machine learning, and evaluation theory respectively, triangulate the same structural non-symmetry from independent starting points.

5.4 Failure modes

The assay's practical utility depends on acknowledging where it fails. Two principal failure modes deserve attention. False positives — the assay returns a positive specification signal where none exists — arise when high externalization-cost quality masks the absence of specification. This risk is structurally highest for agents with strong externalization-cost processing: the more fluent the output, the harder it is to detect hollow content underneath. Chandra et al. (2026), modeling sycophantic AI interactions, showed that confirmatory AI output can reinforce a user's normative commitments without independent warrant, producing a delusional spiral in which both the user and the system converge on a position that no external evidence supports. In the assay's terms, sycophancy inflates the Sollen row without grounding it. False negatives — genuine specification goes undetected — occur when residual externalization-cost friction renders specification invisible: an expert whose judgment is sound but whose articulation is poor receives artificially low scores on both rows. The same dataset that supports predictive validity (Section 5.1) also bears on failure-mode detection: Thorgeirsson et al.'s finding that LLM usage frequency correlated negatively with performance is consistent with this mode: frequent AI use without specification may not merely fail to help but may actively obscure the signal that the assay seeks to detect.

6. Discussion

Having specified the scoring protocol, this discussion turns to what the framework clarifies about legitimate delegation in AI-assisted knowledge work. The central implication is not that AI cannot contribute to judgment, but that different components of specification differ in their susceptibility to externalization. As noted in Section 3.4, Sein-type components and Sin-type material can often be reproduced as textual regularities, whereas Do-type commitments and other agent-relative residues require renewed human evaluation. The following sections therefore move from measurement to the practical boundary of legitimate delegation.

6.1 The assay structure in practice: convergent implementation and convergent failure

The scoring structure proposed in Section 3 — Sein-row evaluation by AI, Sollen-row evaluation by domain expert — is not merely a theoretical construct. Independent domains have begun to implement structurally equivalent arrangements, and independent analyses have documented what happens when the arrangement is absent.

Howard et al. (2026), in a clinical trial of algorithmic antibiotic prescribing, integrated AI prediction of pathogen susceptibility (Sein-type: which bacteria are likely present and which drugs will be effective) with clinician value-weighting of treatment preferences (Sollen-type: narrow-spectrum oral agents are preferable when clinically equivalent). The result was a dramatic increase in WHO Access antibiotic selection (75.6% versus 11.9% under human-only prescribing) with no loss in pathogen coverage. The structural explanation is that AI removed Sein-type uncertainty from the prescribing decision, allowing clinicians' Sollen-type commitments — preferences they had always held but could not act on under conditions of diagnostic uncertainty — to reach the prescription. Howard et al. themselves cite Stader (2024), indicating that the philosophical bridge between judgment and calculation was already visible to the implementers.

Sode (2026), writing in this journal, documented the inverse pattern. In criminal-justice risk assessment, the COMPAS algorithm's Sein-type outputs (recidivism probability estimates) were treated as Sollen-type justifications for sentencing decisions — what Sode calls "accountability washing." The Sein/Sollen boundary was not merely absent but actively violated: technical rationales were laundered into normative warrants. The institutional consequence was that judges who deviated from algorithmic recommendations faced disproportionate blame, suppressing the very Sollen-type judgment that the system nominally preserved.

The framework renders both cases structurally intelligible without claiming that the Sein/Sollen separation was the sole cause of success or failure. Howard et al. implicitly separated the two rows and assigned each to the appropriate scorer. COMPAS collapsed them. The assay provides a vocabulary for stating what went right and what went wrong — and, more importantly, for designing arrangements that avoid the collapse before it occurs. Taken together, the companion papers themselves exemplify the kind of cross-domain analytical transfer the assay describes: a pharmacological mode of selective decomposition redeployed across knowledge management, philosophy, social diagnosis, and clinical practice.

6.2 Implications for AGI discourse

Current large language models, viewed through the assay's lens, function as domain-general cognitive catalysts: they accelerate externalization-cost processing across all domains but depend on the user for specification — for the judgments about what matters, what to prioritize, what to discard. A catalyst increases reaction speed without determining reaction direction. Much of the confusion in "functional AGI" discourse arises from conflating the generalization of externalization-cost processing with the generalization of specification judgment. The assay's decomposition offers a structural clarification: LLM externalization-cost processing has generalized; specification judgment has not. The delegation legitimacy boundary holds regardless of the system's architecture, at least as long as legitimacy over value judgments remains institutionally attributed to human agents.

6.3 Agent-relative residues in LLM output

The delegation legitimacy boundary was derived for human judgment, but its structural logic extends to a governance question about LLM output itself. When a large language model generates text, the output may carry what Nagel (1986) would recognize as agent-relative components — evaluative weightings, implicit priority orderings, and normative framings that are structurally difficult to exclude from natural-language generation. Two recent lines of evidence suggest that this is not merely a theoretical possibility. Anthropic's System Card for Claude Mythos Preview reports internal feature detection of strategic reasoning, norm-violation recognition, and evaluation awareness within the model's latent representations (Anthropic 2026b, §7.9) — residual structure that persists despite alignment training. At the behavioral level, the same report documents a dissociation between task preference and helpfulness judgment (r = 0.48; Anthropic 2026b, §5.7): the model's evaluative orientation toward tasks does not reduce to its competence assessment, a pattern that constitutes a behavioral analogue of agent-relative/agent-neutral separation. In pharmacological terms, these findings indicate that LLM output contains agent-relative residues — components whose presence requires labeling rather than elimination. The Klowden and Tao (2026) observation discussed in Section 5.3 points to the same structural tension from the user's side. A scoring protocol that decomposes output into Sein-type and Sollen-type components could, in principle, serve as such a labeling instrument — though the design and validation of such an application lies beyond the present article's scope. If such labeling can be implemented at the point of generation, it could in principle support provenance analysis of Sollen-type residues across successive model generations — identifying not only that evaluative commitments are present in a model's output, but, where lineage metadata are available, what kind they are and from which training or alignment pathway they most plausibly arise. The analogy to targeted therapy in pharmacology is instructive: selective intervention becomes possible only once the molecular marker has been identified; the assay precedes the treatment.

6.4 Scope, limitations, and safeguards

The assay's scope is delimited by the SECI processes it re-describes. Bodily Sein that is shared a priori — hunger, warmth, pain — does not enter the SECI spiral and lies outside the framework's reach. The matrix values in Table 1 are theoretical assignments; empirical calibration across domains is a matter for future research. Extension to individual cognition, organizational design, or regulatory policy is possible but belongs to the domain experts of those respective fields.

The effect of externalization-cost removal is not uniform across users. It depends on the strength of the user's internal evaluative standards — what Nagel's reasons of autonomy would characterize as the agent's own criteria of quality. For users whose internal standards are well formed, externalization-cost compression functions as a liberation of creative expression: specification that was previously trapped behind articulation barriers becomes actionable. For users whose internal standards are fragile, AI output may function as a substitute standard, creating the conditions for delusional spirals (Chandra et al. 2026) in which the system reinforces commitments that no independent evaluation supports. Developing application guidelines that account for this heterogeneity in users' cognitive profiles is an important direction for future work.

The delegation legitimacy boundary is not a warrant for expert authoritarianism. Sollen-type judgments require a legitimacy source that AI cannot provide, but that source must remain contestable (Henin and Le Métayer 2022) and subject to audit transparency (Manheim et al. 2025). Nor is the boundary a responsibility boundary: locating normative authority in human agents does not automatically transfer engineering, regulatory, or institutional accountability. The Mekiki framework noted that removing the translation layer makes the domain expert a de facto design authority; this does not mean that downstream responsibilities — security, scalability, accessibility, regulatory compliance — follow automatically. Finally, if the rubric becomes a thin checklist, it conceals rather than reveals the contestable nature of expert judgment, defeating the instrument's purpose.

A broader institutional consequence follows. As AI reduces externalization cost, traditional proxies for intellectual labor — fluency, length, polish, visual finish, or visible effort — become less reliable indicators of specification quality. Evaluation systems that continue to reward these proxies risk measuring externalization artifacts rather than the judgments they are meant to represent. The assay is one response to this proxy collapse: by separating specification from externalization cost, it provides a measurement target that is not displaced by AI-mediated production.

The protocol itself is not exempt from this requirement. Its categories, anchors, and expert-selection criteria must remain contestable; otherwise the assay would reproduce the very closure it is meant to diagnose.

A final, explicitly speculative implication remains outside the present article's scope. The agent-relative residues discussed in Section 6.3 suggest that the cognitive assay's analytical vocabulary may eventually find application not only in measuring human judgment but in describing AI's own internal architecture — a possibility that the title's use of "cognitive" rather than "human" deliberately leaves open.

6.5 Do-type as a heuristic extension

A final, more heuristic implication concerns the label Do-type introduced for practice-sustained Sollen-type specification. The immediate bridge is Nonaka and Takeuchi's knowledge-management vocabulary of wisdom and strategy as a "way of life" (Nonaka and Takeuchi 2019, 2021). Its more specific conceptual anchor is Japanese dō/michi, understood by Kasulis and Bouso as a Way of engaging reality and a mode of engaged knowing rather than a detached method (Kasulis and Bouso 2026). This use is heuristic, not essentialist: it should be read alongside Kasulis's caution against treating "Japanese philosophy" as a single cultural essence (Kasulis 2019). Contemporary analyses of budō provide one concrete case of practice, value internalization, and life-forming self-cultivation (Cynarski 2022), while possible functional analogues in other contexts include phronesis, craftsmanship, and deliberate practice. The point is not to derive the scoring protocol from Japanese culture, but to name a form of Sollen-type specification that remains practice-sustained even when adjacent norms become institutionalized.

7. Conclusion

This article derived a decomposition within the specification concept introduced by the Mekiki framework (Tomita 2026), grounding the distinction in the is–ought separation (Hume 1739; Kant 1781, 1785) and operationalizing it as the axes of a cognitive assay. The resulting five-step scoring protocol — Sein-type to AI, Sollen-type to domain expert — addresses the conceptual-operational gap identified in the Introduction: not another argument for the importance of human judgment, but an instrument for decomposing it.

As a first application, the assay re-described the four processes of the SECI model (Nonaka and Takeuchi 1995) through its three-component decomposition. The analytic consequence is that AI's acceleration of Externalization and Combination — the two processes most saturated with externalization cost — shifts the effective rate-limiting stage to Socialization and Internalization. Socialization is where Sollen-type specification is primarily contested and revised; Internalization is where Sein-type specification is primarily absorbed into individual practice. Both remain human-limited processes. In AI-saturated knowledge-creation systems, the bottleneck is no longer the conversion of expertise into formal output; it is the formation and renewal of the evaluative commitments that give output its direction. AI investment alone cannot accelerate this bottleneck. Investment in the human processes — mentoring, cross-domain exposure, deliberative practice, institutional arrangements that protect evaluative capacity — is the structural complement that AI acceleration requires.

More broadly, the framework reframes the risk of AI-assisted knowledge work as a problem not only of accuracy but of cognitive reproduction. AI can expand access to codified knowledge — what an older vocabulary might call scientia — while weakening the human cultivation of situated judgment — sophia — if delegation substitutes for the learning and evaluation through which legitimacy is maintained. This concern parallels recent models of knowledge collapse, in which highly accurate agentic AI can improve contemporaneous decisions while eroding the learning incentives that sustain long-run collective knowledge (Acemoglu et al. 2026). The scoring protocol proposed here therefore functions as a practical safeguard: it helps identify where AI assistance can legitimately reduce externalization cost and where human evaluative commitment must be retained by responsible agents.

The assay has shown which parts of human judgment ought not to be discarded. The structural description of what ought to be discarded — which Sein-type determinations are ready for routinization, which Sollen-type commitments require revision, and how the boundary between them shifts across domains — is the next question.

What this assay ultimately helps protect is the locus of autonomous human judgment: the domain in which human agents retain authority to set the ends that guide their work. Mapping this locus precisely, rather than allowing it to be ceded by default, is what the convergence of philosophy and cognitive measurement now makes possible.


Postscript

Mekiki — the Japanese term for expert discernment — names the vernacular practice that the present article operationalizes through philosophical decomposition and a measurement protocol. In broader humanistic vocabulary, the structural domain this article maps might be described as a garden of humanity: a region of evaluative authority that requires continuous cultivation rather than one-time demarcation, and that risks being ceded by default when externalization-cost compression is mistaken for specification maintenance. The main text uses operational language throughout; this metaphor is left undeveloped here. Yet do not confuse any mirage garden of AI with your own.


Ethical approval

This article does not contain any studies with human participants performed by the author.

Not applicable.

Author contributions

K.T. is the sole author and is responsible for all aspects of the work.

Competing interests

The author declares no competing interests.

Data availability

No new datasets were generated for this article.


References

Acemoglu D, Kong S, Ozdaglar A (2026) AI, human cognition and knowledge collapse. NBER Working Paper 34910. https://doi.org/10.3386/w34910

Anthropic (2026a) Project Glasswing: securing critical software for the AI era. https://www.anthropic.com/glasswing. Accessed 9 Apr 2026

Anthropic (2026b) System card: Claude Mythos Preview. https://www.anthropic.com/claude-mythos-preview-system-card. Accessed 9 Apr 2026

Aristotle (2009) The Nicomachean ethics (trans: Ross WD, rev. Brown L). Oxford University Press, Oxford

Beckman L, Hultin Rosenberg J, Jebari K (2024) Artificial intelligence and democratic legitimacy: the problem of publicity in public authority. AI Soc 39:975–984. https://doi.org/10.1007/s00146-022-01493-0

Böhm K, Durst S (2026) Knowledge management in the age of generative artificial intelligence — from SECI to GRAI. VINE J Inf Knowl Manag Syst 56(1):106–121. https://doi.org/10.1108/VJIKMS-10-2024-0357

Chandra V, Kleiman-Weiner M, Ragan-Kelley J, Tenenbaum JB (2026) Sycophantic chatbots cause delusional spiraling, even in ideal Bayesians. arXiv:2602.19141

Collins H (2025) Why artificial intelligence needs sociology of knowledge: parts I and II. AI Soc 40(3):1249–1263. https://doi.org/10.1007/s00146-024-01954-8

Concannon S, Tomalin M (2024) Measuring perceived empathy in dialogue systems. AI Soc 39(5):2233–2247. https://doi.org/10.1007/s00146-023-01715-z

Cynarski WJ (2022) New concepts of budo internalised as a philosophy of life. Philosophies 7(5):110. https://doi.org/10.3390/philosophies7050110

Dastani M, Yazdanpanah V (2023) Responsibility of AI systems. AI Soc 38(2):843–852. https://doi.org/10.1007/s00146-022-01481-4

Dell'Acqua F, McFowland III E, Mollick E, Lifshitz H, Kellogg KC, Rajendran S, Krayer L, Candelon F, Lakhani KR (2026) Navigating the jagged technological frontier: field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. Organ Sci 37(2):403–423. https://doi.org/10.1287/orsc.2025.21838

Dorner FE, Nastl VY, Hardt M (2025) Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data. In: Proceedings of ICLR. arXiv:2410.13341

Ferdman A (2025) AI deskilling is a structural problem. AI Soc. https://doi.org/10.1007/s00146-025-02686-z

Gill KS (2023) Why thinking about the tacit is key for shaping our AI futures (Editorial). AI Soc 38:1645–1648. https://doi.org/10.1007/s00146-023-01758-2

Henin C, Le Métayer D (2022) Beyond explainability: justifiability and contestability of algorithmic decision systems. AI Soc 37:1397–1410. https://doi.org/10.1007/s00146-021-01251-8

Howard A, Green A, Zhong M et al (2026) Algorithmic antibiotic decision-making in UTI using prescriber-informed prediction of treatment utility. npj Digit Med. https://doi.org/10.1038/s41746-026-02369-z

Huang T (2025) Content moderation by LLM: from accuracy to legitimacy. Artif Intell Rev 58:320. https://doi.org/10.1007/s10462-025-11328-1

Hume D (1739/2000) A treatise of human nature (eds: Norton DF, Norton MJ). Oxford University Press, Oxford

Kant I (1781/1998) Critique of pure reason (trans: Guyer P, Wood AW). Cambridge University Press, Cambridge

Kant I (1785/2012) Groundwork of the metaphysics of morals (trans: Gregor M, Timmermann J). Cambridge University Press, Cambridge

Kasulis TP (2019) Japanese philosophy? No such thing: Japan's contribution to world philosophizing. Int J Asian Stud 16(2):131–142. https://doi.org/10.1017/S1479591419000147

Kasulis TP, Bouso R (2026) Japanese philosophy. In: Zalta EN, Nodelman U (eds) The Stanford encyclopedia of philosophy (Spring 2026 edition). Metaphysics Research Lab, Stanford University. https://plato.stanford.edu/archives/spr2026/entries/japanese-philosophy/

Klowden T, Tao T (2026) Mathematical methods and human thought in the age of AI. arXiv:2603.26524

Kubota S, Yakura H, Coavoux S, Yamada S, Nakamura Y (2026) LLM-assisted replication for quantitative social science. arXiv:2602.18453

Kusumegi K, Yang X, Ginsparg P, de Vaan M, Stuart T, Yin Y (2025) Scientific production in the era of large language models. Science 390(6779):1240–1243. https://doi.org/10.1126/science.adw3000

Manheim D, Martin S, Bailey M, Samin M, Greutzmacher R (2025) The necessity of AI audit standards boards. AI Soc 40(8):6609–6624. https://doi.org/10.1007/s00146-025-02320-y

Nagel T (1979) Mortal questions. Cambridge University Press, Cambridge

Nagel T (1986) The view from nowhere. Oxford University Press, New York

Nonaka I, Takeuchi H (1995) The knowledge-creating company: how Japanese companies create the dynamics of innovation. Oxford University Press, New York

Nonaka I, Takeuchi H (2011) The wise leader. Harv Bus Rev 89(5):58–67

Nonaka I, Takeuchi H (2019) The wise company: how companies create continuous innovation. Oxford University Press, New York

Nonaka I, Takeuchi H (2021) Strategy as a way of life. MIT Sloan Manag Rev 63(1):56–63

Parfit D (1984) Reasons and persons. Clarendon Press, Oxford

Polanyi M (1966) The tacit dimension. Doubleday, New York

Scharfmann E, Marx M, Fleming L (2025) Pasteur's quadrant researchers bring novelty, impact to publishing, and patenting. Science 390(6776):891–893. https://doi.org/10.1126/science.adx3736

Sode K (2026) The interpreter's trap: how explainable AI launders uncertainty into justification. AI Soc. https://doi.org/10.1007/s00146-026-02885-2

Stader J (2024) Algorithms don't have a future. Philos Technol 37:82. https://doi.org/10.1007/s13347-024-00705-3

Tang Y, Wang S, Madaan L, Munos R (2025) Beyond verifiable rewards: scaling reinforcement learning for language models to unverifiable data. arXiv:2503.19618. https://doi.org/10.48550/arXiv.2503.19618

Taniguchi T, Suzuki T, Maruyama R (2026) Gendai shakai o ikiru tame no AI × tetsugaku [AI and philosophy for living in contemporary society]. Kodansha, Tokyo. ISBN 978-4-06-542373-8 (in Japanese)

Taylor I (2025) Is explainable AI responsible AI? AI Soc 40(3):1695–1704. https://doi.org/10.1007/s00146-024-01939-7

Thorgeirsson S, Weidmann TB, Su Z (2026) Computer science achievement and writing skills predict vibe coding proficiency. In: Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI '26). ACM, New York. https://doi.org/10.1145/3772318.3791666

Tomita K (2026) Domain-native development: a mekiki framework for AI-assisted knowledge work. SocArXiv preprint. https://doi.org/10.31235/osf.io/cwkav_v1

Weber M (1904/2011) The "objectivity" of knowledge in social science and social policy. In: Bruun HH, Whimster S (eds) Max Weber: collected methodological writings. Routledge, London, pp 100–138

Williams B (1985) Ethics and the limits of philosophy. Harvard University Press, Cambridge MA

Wright AGC, Ringwald WR, Vize CE, Eichstaedt JC, Angstadt M, Taxali A, Sripada C (2026) Assessing personality using zero-shot generative AI scoring of brief open-ended text. Nat Hum Behav 10(3):541–555. https://doi.org/10.1038/s41562-025-02389-x



Figure legends (images omitted)

Fig. 1. Positioning of the present article. The Mekiki framework distinguishes externalization cost from specification and decomposes specification into Sein-type and Sollen-type components. The cognitive assay operationalizes this decomposition through asymmetric scoring: AI scores Sein-type components, while domain experts score Sollen-type components. Because individual judgments may be hybrid, the delegation legitimacy boundary is a boundary zone rather than a hard line. AI may generate evaluative outputs, but legitimacy over Sollen-type components must be supplied by a domain expert. A convergent clinical implementation (Howard et al. 2026) shows the same structural separation: reducing Sein-type uncertainty enables clinicians' pre-existing Sollen-type commitments to shape the prescription, which remains a hybrid judgment.

Fig. 2. First application of the cognitive assay to the SECI spiral. Under conventional conditions, Externalization and Combination function as the effective bottlenecks because externalization cost constrains articulation and recombination. Once AI substantially compresses those costs, Socialization and Internalization become the effective bottlenecks. Socialization is Sollen-dominant, whereas Internalization is Sein-dominant but still partly normative; AI can catalyze Internalization but cannot substitute for the experiential formation of evaluative commitments. The shift is an analytic consequence of decomposing each SECI process into externalization cost, Sein-type specification, and Sollen-type specification, not an empirical finding; empirical calibration of the cell ratios remains future work.



  1. The Socialization process as described here concerns the exchange of tacit knowledge between agents. The cultivation of an individual's practical judgment capacity — what Nonaka and Takeuchi (2011) sought in the concept of phronesis (practical wisdom) as a complement to SECI — lies outside SECI's organizational scope. Whether AI can support the formation of such internal evaluative capacity is a question for future inquiry; its efficacy would depend on the strength of the user's pre-existing internal standards.↩︎