# Paper T1 — Domain-Native Development: A Mekiki Framework for AI-Assisted Knowledge Work

**Preprint version:** 1  
**Canonical DOI:** https://doi.org/10.31235/osf.io/cwkav_v1

**Kengo Tomita**
Institute of Technology, Shimizu Corporation


---

## Abstract

When experienced professionals retire, organisations lose not their documented procedures but the accumulated judgement that determines how those procedures are applied. This paper proposes a two-component framework—termed the *mekiki* framework, after the Japanese concept of expert discernment (目利き)— for understanding how artificial intelligence (AI) makes this otherwise invisible knowledge structure observable. The framework identifies two conditions governing whether tacit knowledge can be successfully externalised into formal artefacts—software, documents, analytical tools: *externalisation cost* (the technical barriers to articulating and implementing knowledge) and *specification cost* (the barrier constituted by the domain expertise required to determine what should be built, how quality should be judged, and what should be excluded). Drawing on a revelatory case study in which a domain expert with no programming experience developed and publicly deployed a web application in ten working days using AI tools, the study shows that externalisation cost drops dramatically under AI assistance while specification cost persists unchanged. A domain-ablation experiment, in which two independent AI systems produced technically competent but domain-inappropriate applications from the same dataset without expert guidance, provides supporting evidence for the separability of the two components. Grounded in the SECI knowledge management tradition, the framework offers organisations a structural vocabulary for identifying which knowledge components are at risk of irreversible loss through generational transition—and which can be preserved through AI-mediated dialogue while retiring experts remain available.

---

## Keywords

domain-native development, mekiki framework, tacit knowledge, specification cost, externalisation cost, AI-assisted knowledge work, SECI model

---

## 1. Introduction

In knowledge-intensive industries, the retirement of senior practitioners creates a distinctive form of organisational loss. What disappears is not procedural knowledge—which can be documented in manuals and standard operating procedures—but the accumulated judgement that determines how procedures are applied, when they should be overridden, and what problems they fail to anticipate. In the construction industry, where project-specific conditions vary enormously and technical decisions cascade through decades of a building's service life, this form of expertise loss is particularly acute. Japan's construction sector, facing concurrent demographic contraction and an aging technical workforce, represents one of the most severe instances of this challenge globally.

Three researchers at a single corporate research institute illustrate the problem in structural terms.

The first, a senior specialist now retired, accumulated thirty years of expertise in materials evaluation for indoor environmental quality. His published reports and test protocols survive, but the judgement underlying them—which anomalous readings warrant investigation, which client specifications contain hidden contradictions, which measurement conditions produce misleading results—resided in professional intuition that his documentation captured only partially. His knowledge was extensive but largely tacit.

The second, a junior researcher who inherited the senior specialist's portfolio, has access to AI-assisted analysis tools and a comprehensive institutional archive. Yet the translation of this inheritance into new analytical instruments has proven difficult. The bottleneck is not access to information or computational tools—it is the *judgement* required to determine what the tools should do. The information exists; the capacity to specify its application does not.

The third, the present author, works in an adjacent technical domain. Trained in pharmacology (PhD in Medicine) with additional certification in electrical engineering, the author used AI tools to develop and publicly deploy an EV charging infrastructure application in ten working days as a personal project motivated by domain interest—a task that would conventionally require a software development team and months of iteration. The enabling condition was not programming skill (which the author lacked entirely) but deep domain knowledge in the application's subject matter, combined with AI systems that could translate that knowledge into functioning code.

These three individuals shared the same institutional context: identical organisational position, access to AI tools, and comparable educational attainment. What differed was the relationship between their domain expertise and the task at hand. The senior specialist possessed the relevant expertise but lacked tools to externalise it directly. The junior researcher possessed the tools but lacked the expertise to direct them. The author possessed both domain expertise and AI tools, in a domain where the two could be coupled without intermediaries—and produced a working product in days.

This pattern—in which domain knowledge, not technical implementation skill, constitutes the binding constraint on AI-assisted production—has been observed across multiple fields but has not been systematically characterised. Structurally similar decompositions have appeared independently: prediction versus judgement in economics (Agrawal, Gans & Goldfarb, 2018), envisioning gaps in cognitive science (Subramonyam et al., 2024), specification-process-evaluation alignment in interaction design (Terry et al., 2023), and agent-augmented SECI modes in knowledge management (Böhm & Durst, 2026). That multiple disciplines have converged on analogous two-component structures suggests an underlying phenomenon awaiting unified characterisation.

This paper provides that characterisation. We define *domain-native development* (DND) as the direct production of functional artefacts by domain experts using AI as a translation medium, without requiring intermediary professionals (software engineers, technical writers, data architects) to convert domain knowledge into formal implementations. We propose a two-component framework—*externalisation cost* and *specification cost*—that identifies two distinct conditions governing whether tacit knowledge can be successfully converted into working artefacts. Externalisation cost comprises the technical barriers to articulating and implementing knowledge; specification cost comprises the barrier arising from the domain expertise required to determine *what* should be built, *how* quality should be judged, and *what* should be excluded. The framework is grounded in the SECI knowledge management tradition (Nonaka & Takeuchi, 1995), illustrated through a single revelatory case study with domain-ablation experiment, and carries safety implications drawn from the author's pharmacological training.

The framework's social purpose is immediate. The window during which retiring domain experts can externalise their specification capacity through AI-mediated dialogue is finite and closing. A structural vocabulary for this process—one that makes the distinction between transmissible and non-transmissible knowledge components organisationally negotiable—is a precondition for effective intervention.

---

## 2. Conceptual Framework

### 2.1 Two-component decomposition

When domain knowledge is converted into a formal artefact—a software application, a technical document, an analytical tool—two categories of cost are incurred. We term these *externalisation cost* and *specification cost*.

**Externalisation cost** (also referred to as communication cost in practical contexts) denotes the technical barriers to converting tacit knowledge into explicit, formal representations. This includes the effort of articulating domain concepts in language that implementation tools can process, the iterative correction cycles that arise from imprecise articulation, and the translation losses inherent in any conversion from one representational medium to another. The term follows Nonaka's (1994) use of *externalisation* to describe the conversion of tacit knowledge into explicit forms within the SECI (Socialisation–Externalisation–Combination–Internalisation) model of organisational knowledge creation.

**Specification cost** denotes the barrier arising from the domain expertise required to determine what should be built. This encompasses priority setting (which features matter and which are noise), quality criteria (what constitutes an acceptable output), information filtering (what should be deliberately excluded), and organisational context sensitivity (what constraints apply in this particular deployment). Evaluation criteria are included within specification cost rather than treated as a separate component because, in AI-assisted workflows, the capacity to judge output quality draws on the same domain knowledge as the capacity to specify what should be produced; both require knowing what "good" looks like in a given domain. Specification cost is not reducible to requirements elicitation in the software engineering sense; it includes judgements that the domain expert may not be able to articulate as formal requirements but can recognise and evaluate when presented with candidate implementations. Specification cost should be distinguished from domain expertise itself: domain expertise is a resource that the practitioner possesses; specification cost is the degree to which a task demands domain expertise.

Under conventional conditions—human-to-human knowledge transfer, manual document authoring, professional intermediaries—these two cost components are entangled and inseparable.

We term this two-component decomposition the *mekiki* framework, adopting the Japanese term for expert discernment (目利き)—the capacity to see what others cannot see in a domain. Japanese knowledge management concepts have entered international discourse through two distinct routes: practice-driven terms such as *kaizen* and *gemba*, which emerged from Toyota's production system and spread through industrial adoption before being codified in management literature (Imai, 1986; Womack, Jones & Roos, 1990); and theory-driven terms such as *ba* (Nonaka & Konno, 1998) and *SECI* (Nonaka & Takeuchi, 1995), which were introduced as formal concepts in academic publications. *Mekiki* bridges these two lineages: it originates as a vernacular term from Japanese professional practice, but enters international discourse through academic formalisation. Where these earlier concepts addressed the *process* of knowledge creation, the mekiki framework addresses the *cost structure* of knowledge conversion under AI assistance.

The act of specifying (deciding what to build) and the act of externalising (articulating that decision for others to implement) co-occur in every design conversation, every requirements document, and every mentoring relationship. Figure 1 situates this two-component decomposition within the SECI framework.

**Intellectual lineage.** The framework operates within a lineage that runs from Polanyi through Nonaka to the present analysis. Polanyi (1966) established that knowledge contains a dimension that resists articulation: "we can know more than we can tell." Nonaka and Takeuchi (1995) operationalised this insight by treating tacit knowledge as convertible through externalisation, deliberately bracketing Polanyi's deeper claim about inarticulate knowledge for practical purposes. Their later work extended tacit knowledge toward "Wisdom"—knowledge inseparable from lived experience and ethical judgement (Nonaka & Takeuchi, 2019)—though they did not possess an experimental system for separating Wisdom from more readily externalised forms of tacit knowledge. The present framework identifies, within Nonaka's Externalisation phase, two conditions that determine whether externalisation produces domain-appropriate outputs: a component reducible by AI (externalisation cost) and a component (specification cost) whose deeper layers—grounded in private networks, accumulated life experience, and the finite embodied perspective of the knower—structurally correspond to the dimension Polanyi bracketed and to what Nonaka and Takeuchi later termed Wisdom (see Section 5.6, Future directions).

### 2.2 The distillation model of authorship

The relationship between human expertise and AI processing can be modelled as a distillation process. The raw material is collective knowledge—the accumulated information, conventions, prior art, and community practices available within a domain. The apparatus is the AI system, which processes this material while substantially attenuating the interpersonal biases (hierarchy, ego protection, political calculation, social desirability) that characterise human-to-human knowledge transfer.^[The claim is not that AI is bias-free. RLHF training residuals, model personification tendencies, and implicit evaluator assumptions persist. The claim is narrower: the specific interpersonal confounders that make externalisation cost and specification cost inseparable in human dyads—face-saving, status negotiation, evaluation anxiety—are selectively reduced.] The distillation conditions are the human expert's specification: what to retain, what to discard, what quality criteria to enforce, what audience to address.

Under this model, authorship resides in the setting of distillation conditions, not in the operation of the apparatus. The same raw material processed through the same apparatus under different distillation conditions yields different products—a property directly observable when different domain experts use the same AI system to produce artefacts in their respective areas. Conversely, identical distillation conditions applied through different apparatus (different AI models) produce outputs that converge on the same structural features, as the ablation experiment in Section 4 demonstrates.

The model also accounts for iterative refinement. Distillation is rarely a single pass; the expert evaluates intermediate products, adjusts conditions, and re-processes. This multi-pass structure is visible in the dialogue logs deposited with this paper, where specification evolves through progressive refinement rather than appearing fully formed. Figure 2 illustrates this distillation model.

### 2.3 Positioning among existing frameworks

The externalisation cost / specification cost distinction is not the first two-component decomposition proposed for AI-era knowledge work. Table 1 situates the present framework among four independent contributions that address structurally similar phenomena.

**Table 1: Existing two-component frameworks and their relationship to the present model**

| Discipline | Decomposition | Source | Relationship to Ext. cost / Spec. cost |
|---|---|---|---|
| Economics (decision theory) | Prediction / Judgement | Agrawal, Gans & Goldfarb (2018) | Operates at the level of decisions under uncertainty. Spec. cost subsumes judgement but extends to priority setting, quality criteria, and information filtering—dimensions visible in the ablation results (Section 4) where AI produced functional applications lacking not predictions but *design judgements* |
| Cognitive science (HCI) | Gulf of execution / Gulf of envisioning | Subramonyam et al. (2024) | Three gaps (capability, instruction, intentionality) map onto externalisation cost. Specification cost—knowing what the artefact should be—is presupposed as a given goal, whereas the present framework treats it as the primary variable |
| Interaction design | Specification / Process / Evaluation alignment | Terry et al. (2023) | Three-part taxonomy compresses to two components: process alignment and evaluation mechanics map onto externalisation cost; specification and evaluation *criteria* constitute specification cost. Compression from descriptive taxonomy to economic framework enables directional predictions |
| Knowledge management (SECI) | GRAI framework (8 interaction fields) | Böhm & Durst (2026) | Adds a human/machine agent dimension to SECI's four modes. The Ext./Spec. distinction is orthogonal: it applies *within* any SECI mode and any GRAI interaction field, providing an analytical layer that complements their descriptive mapping |

These independent observations—from economics, cognitive science, interaction design, and knowledge management—describe structurally similar phenomena from different disciplinary vantage points. The externalisation cost / specification cost framework provides a parsimonious two-component decomposition that (a) subsumes or is orthogonal to each existing model, (b) connects to an established knowledge management tradition via Nonaka's SECI, and (c) generates directional predictions illustrated through domain-ablation experiment (Section 4). The structural gap extends to AI systems design itself: Tomašev, Franklin, and Osindero (2026) proposed a comprehensive delegation protocol for AI agent systems including task decomposition, monitoring, and verification mechanisms, yet identified no principled criterion for determining which task components can be safely delegated and which require human judgement—precisely the separation that the externalisation cost / specification cost distinction provides.

The framework generates a specific directional prediction: in domains where specification cost is high, AI deployment without domain expert involvement will fail to produce domain-appropriate outputs, regardless of the AI system's externalisation capabilities. Section 4 provides supporting evidence for this prediction.^[The pattern of independent discovery across disciplines is itself informative. Multiple discovery typically indicates that a phenomenon has become observable through shared environmental change—in this case, the widespread availability of generative AI—rather than through the insight of any single researcher (Merton, 1961; Lamb & Easton, 1984).]

### 2.4 Applicability conditions

The substantial reduction in externalisation cost observed in domain-native development applies primarily to individual or small-team production in exploratory phases. Organisational coordination costs—regulatory compliance, multi-stakeholder approval, accessibility auditing, security review—constitute a separate cost category not addressed by the externalisation/specification distinction. Similarly, the framework addresses *technical* externalisation costs—those arising from format conversion, language translation, and procedural compliance. Social and psychological barriers to knowledge sharing (cf. Edmondson, 1999), while relevant to organisational knowledge flows, constitute a distinct phenomenon outside the present framework's scope—though their role as an identification condition is examined in Section 5.1. The framework does not predict that all knowledge work will be transformed equally.

The appropriate unit of analysis is not the task but the *component ratio within a task*. A routine compliance report (high externalisation ratio, modest specification ratio) and a novel product concept (low externalisation ratio, high specification ratio) are both "documents," but their cost compositions differ fundamentally. The framework predicts that AI assistance transforms the former far more than the latter—a prediction consistent with widespread observations of differential AI impact across knowledge work categories.

---

## 3. Case Study

The following case is positioned as a revelatory case study (Yin, 2018)—selected not for prevalence but for its capacity to make observable a phenomenon that the framework predicts but that has not been empirically documented at the level of individual design decisions.

A researcher with a doctoral background in pharmacology and no prior programming experience developed and publicly deployed a progressive web application (PWA) for searching electric vehicle (EV) fast-charging infrastructure across Japanese expressways. The application comprises 32 source files, a bespoke dataset of 583 records (one row per charger model at each directional location, derived from 662 individual charging ports across 461 directional locations at 255 physical sites) with 22 technical fields each, and supports bilingual (Japanese/English) operation. Development spanned ten working days using two AI tools in a deliberate division of labour: an AI chat interface for design discussion and review, and an AI coding agent for implementation. Community reception on social media within 48 hours of public release confirmed practical utility rather than novelty alone.

Time allocation across the development process proved revealing. Actual AI code generation accounted for approximately 5% of total working time. Design discussions, specification review, and iterative refinement consumed roughly 40%. The remainder comprised environment setup (20%), deployment and testing (25%), and breaks (10%). In short, externalisation cost was compressed to a fraction of the total effort; the dominant activity throughout was specification—deciding what the application should and should not do.

### 3.1 Specification cost as the determinant of design

Four categories of design decisions illustrate how domain knowledge, rather than prompt engineering skill, governed application quality.

**Current-based data architecture.** The dataset was structured around voltage and amperage rather than the kilowatt ratings typically displayed by charging networks. This decision derives from the physics of the dominant fast-charging protocol in Japan:^[CHAdeMO, the protocol used by the majority of Japanese fast-charging stations, operates under current control. The charger adjusts output current while voltage tracks the vehicle battery's state of charge.] chargers operate under current control, with voltage tracking the vehicle battery's state of charge. Kilowatt output is therefore a *result*, not a control variable. The dataset accordingly records maximum voltage, maximum amperage, and sustained amperage as separate fields. This architecture enables a vehicle-specific output calculation—min(vehicle voltage, charger voltage) × min(vehicle current, charger current)—that no existing EV information service provides. The formula is self-evident once the underlying physics is understood, yet it remains absent from publicly available charging guides, constituting tacit knowledge that the author's combined electrical engineering certification and EV ownership made available. In the domain-ablation experiment (Section 4), neither AI model operating without this domain knowledge included current-based fields.

**Conditional display suppression.** When a user configures their vehicle profile, the application suppresses boost-mode indicators for chargers whose boost current falls at or below the vehicle's maximum charging current. For instance, a vehicle limited to 200 A receives no benefit from a charger's 350 A boost capability; displaying this information would be misleading. This judgement combines physical understanding (current-limited charging behaviour) with information design principles (progressive disclosure—hiding information irrelevant to the specific user). The suppression logic appeared in neither ablation output.

**Network operator discrimination.** Japanese expressway charging stations are managed by multiple, non-interchangeable networks. Two stations at the same service area may belong to different operators with incompatible authentication systems. The application records operator identity as a distinct field and displays non-default operators with visible badges. In the ablation experiment, one AI model misinterpreted the operator management flag as an "immediate availability" indicator, producing a structurally incorrect data model.

**Non-adoption decisions.** Perhaps the strongest evidence for specification cost lies in what the application deliberately excludes: real-time availability (unreliable API data that could misdirect users), charging time predictions (nonlinear charging curves that vary with temperature, state of charge, and battery degradation), distance calculations (expressway network topology makes simple distance metrics misleading), and tariff information (rates vary by card type, plan, and time of day). Each exclusion reflects domain knowledge about what would harm rather than help users. In the ablation experiment, both AI models exhibited "total field filtration"—converting every available data field into a user-facing filter, unable to judge which information to withhold.

These design decisions were not products of prompt engineering technique. Each originated from the author's domain knowledge—electrical engineering fundamentals, vehicle ownership experience, and understanding of user behaviour at charging stations. They instantiate specification cost as defined in Section 2.

### 3.2 Reversal of the translation direction

Conventional knowledge-to-product workflows follow a linear path: the domain expert writes natural-language requirements, which a software engineer interprets and implements. Ambiguities in the requirements document generate iterative clarification cycles—a primary source of externalisation cost.

The development process observed in this case study inverted this sequence. The author generated working prototypes through AI dialogue, producing functional code that served as the specification itself. When a professional engineer later needed to rebuild components for production quality, the prototype—not a written requirements document—became the reference artefact. A working prototype is structurally less ambiguous than a natural-language specification: behavioural questions can be answered by running the code rather than interpreting text.

This reversal was most evident in a specification-driven development episode. The author prepared a detailed specification document, a visual prototype, and a machine-readable instruction file before engaging the coding agent. The search interface was implemented in eight minutes with zero clarification requests. By contrast, earlier sessions conducted without such preparation required repeated correction cycles as the AI agent guessed at unstated requirements. The contrast demonstrates that externalisation cost reduction is not automatic; it depends on the quality of specification provided.

### 3.3 Exploratory value of reduced externalisation cost

The primary value of reduced externalisation cost may not be the efficiency of producing finished products, but the increased number of exploratory attempts that become feasible when creation cost is low.

The present case study provides direct evidence. The application's specification document records twelve features that were explicitly proposed and rejected during development: a route-distance column (rejected because highway interchange connections introduce measurement ambiguity), a regional filtering system (rejected because certain expressways span multiple administrative jurisdictions, breaking any area-based classification), a power-sharing column (rejected because two-port chargers exhibit variable output distribution that cannot be represented in a single field), and nine additional items, each rejected on domain-specific grounds. In a conventional outsourcing arrangement, most of these features would never have been proposed—and those that were proposed would have incurred significant cost to evaluate and discard.

A design-level pivot illustrates the same dynamic at larger scale. Three complete landing page designs—dark-minimal, clean-light, and editorial—were generated and evaluated within a single session. The editorial approach was selected based on the author's judgement that technical content requires visual authority; the other two were discarded immediately. In traditional outsourcing, commissioning three complete design alternatives and discarding two would represent a substantial cost penalty. Under near-zero externalisation cost, such early-stage exploration becomes routine.

The critical observation is that each rejection was itself an exercise of domain expertise—specification cost applied to *evaluation* rather than creation. The twelve rejected features and two discarded designs represent specification cost in its evaluative dimension: the expert's judgement about what *not* to build.

This dynamic reverses when externalisation cost reduction is applied to organisational automation rather than individual exploration. Individual exploration operates under low failure cost (the explorer bears the consequences); organisational automation operates under high failure cost (errors propagate through systems and users). This asymmetry is examined in Section 5.3.

### 3.4 Two identification points

The case study yields two observations that the Discussion will generalise:

1. **Prototype-as-specification** constitutes a structural mechanism for reducing externalisation cost. The domain expert's output is no longer a natural-language description to be interpreted, but a functioning artefact to be refined.

2. **Increased exploratory attempts** represent a second-order effect of externalisation cost reduction. When creating and discarding prototypes becomes inexpensive, the bottleneck shifts entirely to the domain expert's judgement about what is worth pursuing—specification cost in its evaluative dimension.

---

## 4. Methods

The preceding case narrative and the Discussion that follows draw on three categories of evidence: the dialogue logs generated during AI-assisted development, the outputs of a domain-ablation experiment, and the published dataset. This section describes each.

### 4.1 Dialogue logs as cognitive records

The methodological contribution of this study lies not in the application itself but in the dialogue logs generated during its development. These logs constitute a complete record of the author's decision-making process: which design options were considered, which were adopted and why, and—critically—which were rejected and on what grounds. Rejection rationale is a category of knowledge that typically survives in neither technical manuals nor project reports.

This record shares analytical kinship with Think Aloud Protocols (TAP; Ericsson & Simon, 1993), but with a significant methodological advantage: it is non-intrusive. TAP imposes a dual-task burden by asking participants to verbalise while performing; AI-assisted development naturally produces verbalisation as its primary mode of operation. The dialogue log is not a secondary artefact overlaid on the work process—it *is* the work process. This property also mitigates the primary risk of the author-as-subject design: because decisions were recorded at the moment of occurrence rather than reconstructed retrospectively, the logs constrain post hoc rationalisation in a way that interview-based or memoir-based methods cannot.

A technical precondition for this approach is the context window capacity of current large language models (approximately 200,000 tokens at the time of this study), which permits extended multi-turn dialogues that accumulate project-specific context. This accumulated context creates a positive feedback loop: as the AI system's representation of the project improves, the author's specification becomes more efficient, accelerating knowledge extraction. This technical parameter should be noted as a reproducibility condition.

### 4.2 Domain-ablation experiment

To illustrate the framework's prediction that specification cost cannot be substituted by externalisation cost alone, a domain-ablation experiment was conducted. The experiment is designed not as a benchmark but as a counterfactual illustration: it asks what happens when externalisation capability is present but specification is absent. The same raw dataset—662 charging port records derived from the author's structured CSV but converted to JSON with field names replaced by the network operator's abbreviated API labels (e.g., `im`, `cap`, `typ`), deliberately reducing human readability to prevent domain knowledge leakage through column naming—was provided to two commercially available AI systems: Claude Code (Anthropic, Claude Opus 4.6) and ChatGPT 5.2 (OpenAI, Pro mode). Each received a single generic prompt requesting creation of an EV charging station search application with filtering and sorting capabilities, with design left to the AI's discretion. No follow-up instructions were given. Neither model posed domain-specific queries; the only author interactions during the Claude Code session were technical file-creation approvals, which were set to auto-approve after the first prompt. This one-shot, zero-interaction design prevents contamination from the author's iterative specification.

A critical methodological choice was using the raw API data (JSON with abbreviated field names such as `im`, `cap`, `typ`) rather than the author's structured CSV dataset. The CSV column names—`max_amps`, `sustained_amps`, `boost_minutes`—themselves encode domain knowledge that would partially compensate for specification cost absence.

Both models produced functional applications rapidly: Claude Code generated a single HTML file (353 lines) in under two minutes; ChatGPT produced a single HTML file (1,145 lines, with JSON data embedded inline) after approximately 35 minutes of processing. Both outputs featured competent user interfaces with filtering and sorting capabilities, confirming full externalisation cost capability. Table 2 catalogues the convergent deficiency patterns across three analytical layers.


**Table 2: Convergent deficiency patterns in domain-ablation experiment**

| Analytical layer | Specification-absent output | Specification-present design (author) |
|---|---|---|
| **Data interpretation** | | |
| Physical architecture | Kilowatt ratings only; no current-based data model | Voltage/amperage separation enabling vehicle-specific output calculation (V × A) |
| Temporal resolution | No boost/sustained current distinction | Separate fields for maximum amperage (boost) and sustained amperage |
| Data hierarchy | Flat presentation of all 662 ports (Claude) / Direction-level aggregation into 465 groups without site-level consolidation (ChatGPT) | Three-level hierarchy: 662 ports → 461 locations → 255 sites |
| Operator semantics | Management flag misinterpreted as availability indicator^a^ | Operator identity as distinct field with network-specific badge display |
| **UI/UX** | | |
| Output personalisation | No vehicle-specific calculation | min(vehicle V, charger V) × min(vehicle A, charger A) per configured profile |
| Conditional logic | No display suppression based on user context | Boost indicators hidden when charger boost ≤ vehicle max current |
| Data normalisation | Direction field ignored entirely (Claude) / Eight directional labels retained without normalisation (ChatGPT)^b^ | Normalised to two physical directions per location |
| **Information architecture** | | |
| Navigation model | No route-planning orientation | Route-based (expressway → service area → charger) |
| Curation | No recommendation or prioritisation logic | Operator badges, output ranking, contextual recommendations |
| Filter design | Minimal filters with most fields ignored (Claude) / Extensive filters including single-value fields (ChatGPT); neither selected filters based on user relevance | Selective filtering; irrelevant or misleading fields deliberately excluded |

^a^ One model interpreted the `isManaged` boolean as "immediate availability," producing a structurally incorrect data model.
^b^ One model ignored the direction field entirely; the other retained eight directional labels without normalising to two physical directions per location.

Both models independently exhibited structurally convergent deficiency patterns across all three layers, though with variation in degree. At the *data interpretation* layer: absence of current-based architecture (kilowatt ratings only), no boost/sustained current separation, flat presentation of all 662 ports without grouping in one model and direction-level aggregation (465 groups) without site-level consolidation in the other (versus the author's three-level hierarchy: 662 ports → 461 locations → 255 sites), and misinterpretation of operator management fields. At the *UI/UX* layer: no vehicle-specific output calculation, no conditional display logic, and failure to normalise directional labels (one model ignored the direction field entirely while the other retained eight directional labels, neither normalising to the two physical directions per location used in the author's design). At the *information architecture* layer: no route-planning orientation, no recommendation logic, and—most tellingly—divergent but equally specification-absent filter strategies—one model ignored most fields while the other converted all available fields into filters including one with a single possible value—with neither selecting filters based on domain-informed user relevance.

Each row in Table 2 represents a design decision requiring domain knowledge—electrical engineering fundamentals, charging protocol physics, user behaviour at highway service areas—that no amount of prompt engineering sophistication could substitute. The convergent failure pattern across two independent models operating on identical data is consistent with the framework's prediction: when specification cost is absent, externalisation cost reduction alone produces artefacts that are technically competent but domain-inappropriate. One model's interpretation of the `isManaged` field as "immediate availability" exemplifies the pattern precisely—externalisation capacity (converting a field into a functional UI filter) operated flawlessly, but specification (knowing what the field *means* in the charging network context) was absent.

### 4.3 Data availability

Three data layers are deposited on figshare (Nature Portfolio's recommended repository): the charging infrastructure dataset (583 × 21 CSV; one field—installation date, collected through independent research—is withheld from the public release), the ablation experiment input data (662-record JSON with abbreviated field names, as provided to both AI systems) together with outputs and session logs, and full dialogue logs in the original Japanese (machine-analysable by contemporary AI systems, rendering language a non-barrier to verification). Dialogue logs have been reviewed to remove references to third-party individuals and proprietary organisational information; redacted passages are marked with [REDACTED] tags to maintain analytical traceability.^[To maintain double-anonymous review integrity, data are submitted as supplementary materials during review. The figshare deposit will be made publicly accessible upon acceptance, with the DOI reserved in advance.]

### 4.4 AI use disclosure

In accordance with Springer Nature's policy on AI-assisted writing, the following disclosure is provided. Three AI systems were used throughout this research: an AI chat interface (Claude, Anthropic) for design discussion, conceptual development, and manuscript drafting; an AI coding agent (Claude Code, Anthropic) for software implementation; and an additional AI system (ChatGPT, OpenAI) for the domain-ablation experiment (Pro mode, version 5.2) and independent manuscript review (Pro mode, five review cycles). The AI chat interface also served as the primary data collection instrument, generating the dialogue logs analysed in this study.

In all cases, the author provided domain knowledge, made all design and analytical decisions, set evaluation criteria, and bears sole responsibility for the content. The AI systems performed externalisation—converting the author's specifications into formal text, code, and structured arguments. No AI system is listed as an author, consistent with COPE and ICMJE guidelines. The dialogue logs deposited with this paper permit verification of the human–AI division of labour at the level of individual decisions.

---

## 5. Discussion

### 5.1 Theoretical positioning: the visibility thesis

Section 3 demonstrated two identification points: that prototype-as-specification constitutes a structural mechanism for reducing externalisation cost (identification point 1), and that reduced externalisation cost increases the rate of exploratory attempts whose evaluation depends on specification cost (identification point 2). This section examines the theoretical implications of these observations.

The central claim of this paper is that the distinction between externalisation cost and specification cost is not a structure created by AI, but a structure *made visible* by AI. The two components have always been present in knowledge work; they were simply inseparable under prior conditions.

The reasoning follows a quasi-experimental logic. In conventional human-to-human knowledge transfer, multiple confounding variables overlay the translation process: hierarchical relationships between mentor and apprentice, evaluation anxiety, social desirability bias, organisational politics, time pressure, and ego protection. These interpersonal factors are entangled with both the technical difficulty of articulating tacit knowledge (externalisation cost) and the substantive judgements about what knowledge matters (specification cost). Under these conditions, the two cost components cannot be independently observed.

The separation is not unique to AI. Earlier waves of information technology already eliminated narrow segments of externalisation cost—institutional databases that auto-populate grant applications from publication records, for instance, removed the purely clerical effort of transcribing one's own work, a task requiring zero judgement. Yet these reductions targeted only routinised, judgement-free components of externalisation. The residual externalisation cost remained large enough to keep specification cost empirically indistinguishable from the total. What generative AI altered was the *breadth* of externalisation cost reduction—extending from text generation and translation to code synthesis and structural organisation—thereby rendering the persistent specification cost component visible for the first time as a distinct, measurable residual.

AI-mediated dialogue substantially attenuates these confounders. Iteration carries near-zero marginal cost, social penalties for reformulation are minimal, explanation requests carry no status implications, and power asymmetries are substantially reduced in the dyadic AI-human interaction context. Under these attenuated conditions, what remains observable is the differential behaviour of the two cost components: externalisation cost drops dramatically (eight-minute implementation episodes, zero-iteration specification transfer), while specification cost persists unchanged (the design decisions catalogued in Section 3 required the same domain knowledge regardless of the implementation medium).

The domain-ablation experiment (Section 4.2) provides the identification step. Under the specification-absent condition (AI operating without domain knowledge), externalisation capacity alone produced technically competent but domain-inappropriate artefacts. The convergent failure pattern across two independent AI models—exhibiting the same categories of deficiency despite different architectures and different interface choices—suggests that the two components are separable and that they respond differently to AI assistance. This pattern is consistent with the defining prediction of the framework.

This structure is not unique to the present case study. The independent emergence of structurally analogous two-component decompositions across economics (Agrawal et al., 2018), cognitive science (Subramonyam et al., 2024), interaction design (Terry et al., 2023), and knowledge management (Böhm & Durst, 2026), as documented in Section 2, constitutes convergent evidence for the universality of the underlying structure.

A finer-grained observation follows from the case study data, though it extends beyond the ablation evidence into the author's reflexive analysis. The two-component structure is an ontological claim: externalisation cost and specification cost are always co-present in knowledge work. The simultaneous attenuation described below is an epistemological condition—it enabled observation, not existence. Externalisation cost itself contains at least two subcomponents that respond to different interventions: *interpersonal risk perception*—the perceived social cost of articulating tentative or incomplete knowledge, which psychological safety attenuates (Edmondson, 1999)—and *technical conversion barriers*—the procedural difficulty of translating knowledge into formal representations, which AI attenuates. In the present case, both subcomponents were simultaneously reduced: technical conversion barriers by AI-mediated dialogue, and interpersonal risk perception by the author's low emotional responsiveness profile (Section 5.6). This simultaneous attenuation created the observational conditions under which specification cost could be isolated as a pure residual. The subcomponent distinction suggests that maximising specification cost utilisation may require combined deployment of psychological safety interventions and AI tools, each targeting the externalisation barrier it can address.

### 5.2 Domain knowledge as the substance of specification cost

If externalisation cost is what AI reduces, specification cost is what determines output quality under the new cost structure. The case study's design decisions (Section 3.1) provide direct evidence: each reflects domain knowledge that could not have been generated by prompt engineering technique, however sophisticated.

The unit of analysis for this framework is not the task but the *component ratio within a task*. A routine technical report (high externalisation ratio, low specification ratio) and a novel research proposal (low externalisation ratio, high specification ratio) are both "writing tasks," but their cost compositions differ fundamentally. AI assistance transforms the former far more than the latter.

The portability of specification cost provides additional evidence for its locus in the human contributor. The structured handoff files used to transfer project context between AI sessions—documents that encoded accumulated design decisions and their rationale—reduced externalisation cost in new sessions to near zero from the first exchange. The fact that specification could be extracted, encoded, and successfully transferred suggests that it resides in the human domain expert's judgement rather than in any particular AI system's learned representations.

The framework is not AI-specific. An organisational case illustrates the same principle without AI involvement: an automobile manufacturer's hybrid vehicle displays domain-specific aerodynamic features—quantified drag-reduction indicators presented directly to consumers during driving—that would be unlikely to emerge from a conventional requirements-specification-implementation pipeline without direct involvement of aerodynamic expertise in the design process. The structural explanation is identical: embedding domain expertise eliminates the translation layer, allowing specification to flow directly into the artefact.

The framework further predicts that the domain expertise underlying specification cost, being grounded in pattern recognition and structural judgement rather than domain-specific vocabulary, can operate across disciplinary boundaries once externalisation barriers are removed. Systematic replication by domain experts in other fields constitutes a priority for future research.

### 5.3 Non-functional requirements and the responsibility boundary

Domain-native development reduces externalisation cost for domain-aligned functionality—the features and design decisions that directly reflect the expert's knowledge. It does not address non-functional requirements: security, scalability, accessibility, long-term maintainability, and regulatory compliance. These constitute a separate axis of professional expertise.

The translation layer that DND eliminates served a dual function. It was simultaneously a source of information loss (externalisation cost) and a quality gate staffed by professionals trained in system-level concerns (responsibility boundary). When the translation layer disappears, the domain expert becomes the de facto design authority—a role that carries responsibilities beyond domain correctness.

A reported industrial case illustrates the risk at scale: a major IT services firm deployed a multi-agent AI platform that automated the full software modification lifecycle—from requirements analysis through design, implementation, and integration testing—achieving approximately 100-fold productivity gains on selected cases from a pool of roughly 300 modification tasks (Fujitsu, 2026). When such automation removes human intervention across the entire development lifecycle, the failure containment boundary expands proportionally—a defect in the automation logic could propagate through requirements, design, and implementation without the stage-gate reviews that conventional team-based development provides.

The asymmetry identified in Section 3.3 is structural. Individual exploration (where the domain expert bears the consequences of failure) operates safely under reduced externalisation cost. Organisational deployment (where failures affect users, systems, and regulatory obligations) requires that the responsibility boundary be explicitly reconstructed—through code review, security auditing, or professional engineering oversight—even when the translation layer has been removed.

A broader observation follows. Not all externalisation cost reduction is beneficial.^[An example within the author's development: the AI coding agent flagged a potential security concern (unprotected API endpoint) that the author, lacking security expertise, would not have identified independently. The pattern—AI detection followed by human approval or rejection—suggests a general architecture for responsibility management in DND contexts, where AI systems can identify potential issues within their training distribution but cannot determine whether the issue is material in the specific deployment context.] In established professional communities, high externalisation cost can function as a *deliberate* entry barrier, maintaining quality standards and in-group coordination efficiency. This is not a cultural phenomenon limited to specific national contexts; it is a structural property of any group that has accumulated shared tacit knowledge. Reducing externalisation barriers for outsiders may dilute the specification density that gives the community its productive advantage—a tradeoff that organisational AI adoption strategies must evaluate explicitly.

### 5.4 Methodological contribution of dialogue logs

The dialogue logs generated through AI-assisted development constitute a methodological contribution independent of the present framework. Unlike traditional Think Aloud Protocols (Ericsson & Simon, 1993), which impose a secondary cognitive load on participants, AI dialogue produces verbalised reasoning as a natural byproduct of the primary task. The researcher does not need to "think aloud"—the interaction medium *requires* articulation.

The resulting records capture categories of knowledge that are absent from conventional documentation. Design rationale ("why this architecture rather than alternatives"), rejection rationale ("why this feature was excluded"), and evaluative criteria ("what makes this output acceptable") are preserved in the dialogue but would not appear in a requirements document, technical manual, or project postmortem. For knowledge-intensive organisations facing expertise succession challenges, dialogue logs represent a fundamentally new type of institutional memory.

A self-reinforcing dynamic was observed: as the AI system accumulated project context through extended dialogue, the author's subsequent specifications became more efficient—less explanation was required to convey intent. This profile-accumulation effect constitutes an operationalizable measure of externalisation cost reduction over time.

The transferability of this accumulated context was demonstrated through structured handoff documents. Dialogue-derived project context, when systematically encoded, reduced externalisation cost in new AI sessions to near zero, confirming that the valuable knowledge resided in the specification structure rather than in any particular dialogue history.

### 5.5 Cumulative optimisation risk and educational implications

**Risk.** Extended use of a single AI system creates a risk of cognitive lock-in. Individual AI models exhibit characteristic interaction patterns, and sustained optimisation toward one model's affordances may narrow the user's specification repertoire. The author's practice of distributing tasks across two AI systems (Claude for extended design dialogue, and ChatGPT for independent review) itself constitutes a meta-level specification judgement: knowing which AI system to deploy for which cognitive task requires domain knowledge about the systems' respective strengths—a form of meta-specification.

The mitigation for cumulative optimisation risk is framework literacy—understanding the distinction between externalisation cost and specification cost at a conceptual level. This understanding enables practitioners to recognise when they are optimising externalisation (productive) versus when they are inadvertently outsourcing specification (dangerous). Risk recognition is itself a specification-side judgement that protects the practitioner's evaluative capacity.

**Educational implications.** This connection between risk mitigation and conceptual understanding leads directly to the question of whether specification cost can be taught. While externalisation cost reduction operates primarily through sustained delegation to AI systems, specification cost acquisition presents a harder educational challenge. Drawing on the author's experience as both framework developer and training designer, a three-layer model is proposed:

- *Layer 1: Conceptual awareness.* The learner understands the framework intellectually but cannot apply it. Knowing that specification cost exists does not confer the ability to exercise it.
- *Layer 2: Experiential recognition.* Through domain practice, the learner develops pattern recognition that functions with deliberate effort. Time-intensive but achievable.
- *Layer 3: Integrated operation.* Conceptual knowledge and experiential pattern recognition merge, enabling real-time specification—whether during AI interaction, peer review, or independent design judgement.

Training programmes (Layer 1) do not produce specification competence in isolation, but they function as primers that, in pharmacological terms, *sensitise* the learner to subsequent domain experiences. A training session that introduces the externalisation/specification distinction may not immediately change behaviour but may accelerate the recognition of specification opportunities when they arise in practice.

A practical implication follows for organisational training design. The instruction "teach the veteran to use AI" misidentifies the bottleneck. The veteran already possesses specification capacity; what is needed is a training designer who constructs the interaction context—selecting appropriate prompts, structuring the dialogue sequence, choosing which domain knowledge to surface—so that the veteran's specification flows naturally into the AI system. The training designer operates at a meta-specification level: specifying how specification itself should be elicited.

Observable indicators of successful framework acquisition were identified in the author's own experience: spontaneous application of the two-component distinction in everyday professional conversation, without deliberate reference to the framework. This shift from analytical application to automatic categorisation may serve as a measurable proxy for Layer 3 integration.

The distinction between learner-stage users (who lack specification capacity and must develop it through domain apprenticeship) and expert-stage users (who possess specification capacity and need only externalisation cost reduction) is foundational to educational design. Conflating the two populations—as much current "AI literacy" training does—produces programmes that optimise externalisation for people whose bottleneck is specification. This educational analysis leads directly to the question of what institutional structures currently produce specification capacity—a question taken up in the Conclusion.

### 5.6 Safety considerations and limitations

The effectiveness of AI as a cognitive environment carries direct safety implications. In pharmacological terms: a compound that produces therapeutic effects also carries the potential for adverse effects, and the appropriate response is not prohibition but usage design—dosage, contraindications, and monitoring protocols.

AI dialogue acts on cognitive processes with a directness that conventional tools do not. The attenuation of interpersonal buffers—the same property that enables the quasi-experimental separation of cost components—simultaneously reduces the psychological distance between the user and their own knowledge structures. An analogy to opioid regulation is instructive: unrestricted access led to dependency crises, while excessive restriction denied access to patients with legitimate need. The appropriate response was neither extreme but a structured protocol balancing efficacy and risk.

Empirical grounding for this concern emerged from the author's own experience. Despite scoring in the lowest quartile on a standardised measure of emotional responsiveness, the author observed a graduated increase in emotional engagement as the depth of AI-mediated externalisation progressed—from self-knowledge articulation through interpersonal relationship analysis to intellectual lineage positioning. The a fortiori implication is that users with greater emotional responsiveness may experience earlier and stronger reactions. Detailed data are provided in the supplementary materials.

An external case reinforces this concern from a different angle. Ordak (2023) demonstrated that when ChatGPT was asked to recommend statistical tests for published allergology studies, it produced inappropriate test selections, omitted assumptions necessary for valid application, and returned contradictory recommendations for identical questions posed on separate occasions. This constitutes a case of externalisation cost (executing statistical reasoning) being delegated without specification cost (judging whether a particular test is methodologically appropriate for the data structure at hand). The structural parallel to the ablation experiment is direct: ablation showed specification-absent AI producing domain-inappropriate applications; Ordak showed specification-absent AI producing methodologically unreliable statistical guidance.

A pharmaceutical analogy clarifies a structural gap in current practice. Drug labels are required to disclose contraindications alongside therapeutic indications—specifying not only what a compound can treat but where its use is inappropriate or dangerous. No equivalent disclosure framework currently exists for AI-assisted development tools. The mekiki framework offers a conceptual basis for such disclosure: distinguishing tasks where domain expertise is required (high specification cost) from those that can be safely delegated to AI processing (high externalisation cost ratio). Developing this distinction into operational disclosure standards is a task for the broader community—across technology, medicine, and policy—rather than for any single framework.

**Limitations.** This study has several limitations that bound its claims. (1) The case study is n = 1. It is positioned as a revelatory case (Yin, 2018)—valuable for demonstrating the existence and structure of a phenomenon rather than for establishing prevalence. (2) The author serves as both researcher and subject, introducing potential confirmation bias. Full data disclosure (dialogue logs, code, ablation outputs) mitigates but does not eliminate this concern. The observed separation may have been facilitated by simultaneous attenuation of both externalisation cost subcomponents: technical barriers by AI, and interpersonal risk perception by the author's low emotional responsiveness profile. Replication under alternative conditions—such as psychologically safe team environments (Edmondson, 1999) substituting for individual trait-based attenuation—would strengthen generalisability. (3) Non-functional requirements (security, scalability, accessibility) were not systematically evaluated and remain outside the study's scope. (4) Long-term maintainability of the developed application is untested. (5) Externalisation cost reduction was observed in a single-developer context; organisational settings introduce coordination costs not addressed here. (6) The framework's scope is bounded to domains where specification originates from human expertise; domains where AI systems themselves participate in specification generation (e.g., unsolved mathematical conjectures, novel molecular design) may require a different analytical structure. (7) Cognitive and emotional load was assessed through retrospective self-report calibrated against a standardised baseline; physiological measures (heart rate variability, electrodermal activity) would provide stronger evidence but were not available in this study. Generalisability of the emotional engagement pattern requires replication with standardised instrumentation.

**Future directions.** Six avenues merit investigation: (a) Replication by domain experts in other fields to determine whether "AI-irreducible residue" consistently maps onto specification cost. (b) Internal structure of specification cost—preliminary observations suggest a surface layer (externaliseable through iterative AI dialogue) and a deeper layer (grounded in private networks, life experience, and finite embodied existence) that may correspond to what Polanyi (1966) identified as the fundamentally inarticulate dimension of knowledge and what Nonaka and Takeuchi (2019) termed Wisdom. Specification cost may contain strata that are separable in principle, an architecture whose investigation requires controlled comparison across expertise levels. (c) The relationship between context window capacity expansion and specification portability—whether larger context windows reduce specification cost or merely reduce externalisation cost for specification transfer. (d) Longitudinal observation of threshold effects as AI systems improve in resolution, scope, and calibration. (e) Extension of the two-component decomposition principle to Socialisation (tacit-to-tacit transfer), particularly in domains involving bodily skill transmission (e.g., construction trades, surgical training). A quasi-experimental system analogous to LLMs—capable of selectively removing bodily transmission barriers while preserving embodied judgement—does not yet exist, but emerging technologies (VR, motion capture, robotics) may eventually serve this role, warranting separate investigation. (f) Systematic investigation of externalisation cost substructure. This study identified at least two subcomponents—interpersonal risk perception (Edmondson, 1999) and technical conversion barriers—that respond to different interventions. Whether additional subcomponents exist, and how they interact under varying organisational conditions, remains unexamined.

---

## 6. Conclusion

This paper introduced the mekiki framework, which identifies two conditions governing whether tacit knowledge can be successfully converted into formal artefacts: externalisation cost (the technical barriers to articulating and implementing knowledge, corresponding to Nonaka's Externalisation phase in the SECI model) and specification cost (the barrier constituted by the domain expertise required to determine what should be built, how quality should be judged, and what should be excluded). Using a distillation model of authorship—where collective knowledge serves as the raw material, AI provides a bias-attenuated processing environment, and the human expert sets the distillation conditions—the framework redefines authorship in AI-assisted knowledge work as the exercise of specification.

### The PhD as specification training

The framework permits a structural redescription of doctoral training. The formal definition of a PhD—"an original contribution to knowledge through original research" (QAA, 2020; EUA, 2019)—decomposes naturally: *identifying a structure* is the generation of domain-specific specification; *converting it to formal knowledge* is externalisation in the SECI sense; *achieving the precision required for peer review* is externalisation cost processing at the standards of the academic community.

Doctoral training is, under this analysis, a specification training apparatus. AI is an externalisation cost reduction apparatus. The conjunction of the two constitutes the enabling condition for domain-native development—a conjunction this paper itself instantiates, having been written as a single-authored conceptual contribution in knowledge management by a researcher trained in pharmacology.

A policy implication follows. Discussions of doctoral workforce utilisation often focus on technical skills transfer or interdisciplinary breadth. The framework suggests that the core asset of doctoral training is neither—it is the *capacity to specify*: to identify what matters in a domain, to judge quality against implicit standards, and to determine what should not be built. When externalisation costs approach zero, this capacity becomes the binding constraint on productive AI use.

### Observed diffusion and framework-derived conjectures

The specification/externalisation dynamic is observable across domains at varying stages of AI integration. In competitive strategy games, AI superiority has been established for over two decades; the surviving human contribution is precisely the strategic judgement that the framework terms specification cost—the ability to select among AI-recommended options based on contextual evaluation (Shoresh & Loewenstein, 2025). In professional translation, the shift from creation to post-editing has redistributed labour from externalisation to specification (Flanagan et al., 2025). In software engineering, recent commentary on "vibe coding" (Sarkar & Drosos, 2025) observes that programming expertise is being redistributed toward evaluation—a restatement of the externalisation-to-specification shift in domain-specific vocabulary. In academic and professional writing, the present paper itself instantiates the pattern: a domain expert producing a conceptual contribution in an unfamiliar disciplinary genre through AI-mediated externalisation.

Two conjectures follow from the framework but exceed the present study's evidence. First, in visual illustration, the residual specification cost may be smaller than practitioners assume—technical rendering skill, long classified as specification, may be reclassifiable as externalisation cost as generative models improve. The professional implications require careful analysis that the present framework can structure but not resolve. Second, "high-context" communication barriers are not cultural phenomena but structural properties of *any* group with accumulated shared tacit knowledge. Externalisation cost reduction is therefore not uniformly beneficial; it can erode the specification-dense coordination that gives expert communities their productive advantage. Organisations in industries where knowledge is distributed at the operational level—construction, manufacturing, specialised services—face this tradeoff most acutely.

### Urgency

The framework's practical value is proportional to the speed at which it reaches practitioners. Domain experts whose domains carry high specification cost—typically senior professionals approaching retirement—represent a finite window during which their knowledge can be captured through AI-mediated dialogue logs. The three researchers introduced in Section 1 exemplify this asymmetry: the retiring specialist's specification capacity is irreversibly depleting, while the junior researcher's externalisation cost reduction tools await specification that only the specialist can provide. Each retirement without such capture represents an irreversible loss of specification capacity that no amount of subsequent externalisation cost reduction can compensate. The time structure is asymmetric: externalisation cost will continue to decrease as AI systems improve, but specification cost, once lost through generational transition, cannot be regenerated from data alone.

The urgency is not limited to knowledge preservation. The distillation model described in this paper—in which AI provides a bias-attenuated processing environment while humans set distillation conditions—carries its own time-sensitive implication: the framework for distinguishing what should and should not be delegated to AI must be established before widespread adoption outpaces the institutional capacity to set those conditions.

### Closing

Naming a phenomenon is the precondition for managing it. The distinction between externalisation cost and specification cost has always been present in knowledge work—in the experienced engineer's frustration when a junior colleague's AI-generated report is technically fluent but substantively empty, in the architect's inability to explain why one design "works" and another does not, in the physician's diagnostic judgement that resists algorithmic capture. What was missing was a vocabulary that made this distinction negotiable within organisations, measurable within research, and teachable within educational programmes.

This paper has proposed such a vocabulary, grounded it in the SECI knowledge management tradition, illustrated it through a single revelatory case, and provided supporting evidence for its central prediction through domain-ablation experiment. The framework does not claim to resolve the relationship between human expertise and AI capability. It claims something more modest and more immediately useful: to make the structure of that relationship *visible*—a structure that, as Lovelace (1843) noted of computation itself, has always concerned the boundary between what machines can perform and what only those already acquainted with the domain can provide.

The mekiki framework presented here does not create new capabilities; it makes visible what was always there.

---

## Postscript: A general form of domain-native development

The case study presented in this paper documents a *special form* of domain-native development: one in which specification cost was not pre-existing but was acquired and refined through AI dialogue itself. The framework, the terminology, and the analytical structure emerged during the process. This special form was difficult and time-consuming — ten working days for the application, substantially longer for the conceptual framework — precisely because specification and externalisation were co-developing.

Following submission of this paper, the author conducted three independent writing projects in domains where specification cost had accumulated over approximately ten years each: musical history, community governance, and research-support infrastructure. In each case, the author’s role was limited to confirming and adjusting specification — what to include, what to exclude, what level of detail was appropriate — while AI handled externalisation. All three documents were completed within a single day.

The contrast is direct. Same author, same AI, same period. The only variable was whether specification cost pre-existed. Where it did, production was rapid and straightforward. Where it did not (the present paper), production required iterative co-development over weeks.

This difficulty inversion is itself evidence for the framework’s central claim: externalisation cost and specification cost are separable, and they respond independently to AI assistance. For practitioners reading this preprint: if you possess deep domain experience, the general form of domain-native development is already available to you. The bottleneck was never the tool.

---

## Ethical approval

This article does not contain any studies with human participants performed by any of the authors. The study analyses the author's own AI-generated dialogue logs produced during a software development project; no data from other individuals was collected or analysed. The three individuals described in Section 1 are presented as anonymised structural illustrations of organisational knowledge patterns and were not research participants.

## Informed consent

This article does not contain any studies with human participants performed by any of the authors. As no human participants were involved, informed consent was not applicable.

## Competing interests

The author declares no competing interests. The author is employed by a construction company whose knowledge management challenges are discussed in the Introduction as contextual motivation; the company had no role in the study design, data collection, analysis, or decision to publish.

## Data availability

The datasets generated and analysed during the current study are available in the figshare repository. Three data layers are deposited: (1) the charging infrastructure dataset (583 × 21 CSV; one field withheld from public release), (2) the ablation experiment input data (662-record JSON) together with outputs and session logs, and (3) representative dialogue logs in the original Japanese. During double-anonymous review, data are submitted as supplementary materials; the figshare deposit will be made publicly accessible upon acceptance. All materials are licensed under CC BY-NC 4.0. The datasets generated and analysed during the current study are available in the figshare repository (https://doi.org/10.6084/m9.figshare.31441864).

---

## References

1. **Agrawal, A., Gans, J., & Goldfarb, A. (2018).** *Prediction machines: The simple economics of artificial intelligence.* Harvard Business Review Press.

2. **Böhm, K., & Durst, S. (2026).** Knowledge management in the age of generative artificial intelligence — from SECI to GRAI. *VINE Journal of Information and Knowledge Management Systems, 56*(1), 106–121. https://doi.org/10.1108/VJIKMS-10-2024-0357

3. **Edmondson, A. (1999).** Psychological safety and learning behavior in work teams. *Administrative Science Quarterly, 44*(2), 350–383. https://doi.org/10.2307/2666999

4. **Ericsson, K. A., & Simon, H. A. (1993).** *Protocol analysis: Verbal reports as data* (Rev. ed.). MIT Press.

5. **Flanagan, M., Dam Jensen, H., Bundgaard, K., & Paulsen Christensen, T. (2025).** Technology adoption among Danish translators: Practices, perceptions and prospects. *Revista Tradumàtica*, 23, 111–136. https://doi.org/10.5565/rev/tradumatica.511

6. **Fujitsu. (2026, February 17).** AI-Driven Software Development Platform: Multi-agent AI platform automating the full software modification lifecycle [Press release]. https://global.fujitsu/en-global/pr/news/2026/02/17-01

7. **Imai, M. (1986).** *Kaizen: The key to Japan's competitive success.* McGraw-Hill.

8. **Ordak, M. (2023).** ChatGPT's skills in statistical analysis using the example of allergology: Do we have reason for concern? *Healthcare, 11*(18), 2554. https://doi.org/10.3390/healthcare11182554

9. **Lamb, D., & Easton, S. M. (1984).** *Multiple discovery: The pattern of scientific progress.* Avebury.

10. **Merton, R. K. (1961).** Singletons and multiples in scientific discovery: A chapter in the sociology of science. *Proceedings of the American Philosophical Society, 105*(5), 470–486.

11. **Nonaka, I. (1994).** A dynamic theory of organizational knowledge creation. *Organization Science, 5*(1), 14–37. https://doi.org/10.1287/orsc.5.1.14

12. **Nonaka, I., & Takeuchi, H. (1995).** *The knowledge-creating company: How Japanese companies create the dynamics of innovation.* Oxford University Press.

13. **Nonaka, I., & Konno, N. (1998).** The concept of "Ba": Building a foundation for knowledge creation. *California Management Review, 40*(3), 40–54. https://doi.org/10.2307/41165942

14. **Nonaka, I., & Takeuchi, H. (2019).** *The wise company: How companies create continuous innovation.* Oxford University Press.

15. **Polanyi, M. (1966).** *The tacit dimension.* Doubleday.

16. **QAA (2020).** *Characteristics statement: Doctoral degree.* Quality Assurance Agency for Higher Education. https://www.qaa.ac.uk/the-quality-code/characteristics-statements/characteristics-statement-doctoral-degrees

17. **EUA (2019).** *Doctoral education in Europe today: Approaches and institutional structures.* European University Association. https://eua.eu/resources/publications/809:doctoral-education-in-europe-today-approaches-and-institutional-structures.html

18. **Sarkar, A., & Drosos, I. (2025).** Vibe coding: Programming through conversation with artificial intelligence. In *Proceedings of the 36th Annual Conference of the Psychology of Programming Interest Group (PPIG 2025).* arXiv:2506.23253

19. **Shoresh, D., & Loewenstein, Y. (2025).** Modeling the centaur: Human-machine synergy in sequential decision making. In *Proceedings of the 24th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2025)* (pp. 1941–1949). IFAAMAS. https://doi.org/10.5555/3709347.3743831

20. **Subramonyam, H., Pea, R., Pondoc, C., Agrawala, M., & Seifert, C. (2024).** Bridging the gulf of envisioning: Cognitive challenges in prompt based interactions with LLMs. In *Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI '24).* ACM. https://doi.org/10.1145/3613904.3642754

21. **Terry, M., Kulkarni, C., Wattenberg, M., Dixon, L., & Morris, M. R. (2023).** Interactive AI alignment: Specification, process, and evaluation alignment. arXiv:2311.00710

22. **Tomašev, N., Franklin, M., & Osindero, S. (2026).** Intelligent AI delegation. arXiv:2602.11865

23. **Womack, J. P., Jones, D. T., & Roos, D. (1990).** *The machine that changed the world.* Free Press.

24. **Yin, R. K. (2018).** *Case study research and applications: Design and methods* (6th ed.). SAGE.

---

---

## Figure legends (images omitted)

### Figure 1

**Figure 1. The *mekiki* framework as a pathway model within the SECI model.**

The SECI model (Nonaka & Takeuchi, 1995) describes four modes of knowledge conversion: Socialisation (tacit → tacit), Externalisation (tacit → explicit), Combination (explicit → explicit), and Internalisation (explicit → tacit). Within Externalisation, the *mekiki* framework identifies two conditions that determine whether domain-appropriate outputs are produced. The substrate is *specification* — the domain expertise invested in deciding what should be built, how quality should be judged, and what should be excluded. **Externalisation cost** is the technical barrier to converting that specification into formal representations, and AI selectively reduces this barrier. **Specification cost** is the degree to which the task demands domain expertise. Output quality is therefore dose-dependent rather than binary: more relevant domain expertise yields more domain-appropriate output, while near-zero substrate yields domain-inappropriate output, as illustrated by the domain-ablation experiment. Deeper layers of specification cost structurally correspond to Polanyi's inarticulate dimension and to Nonaka and Takeuchi's (2019) Wisdom, beyond the present framework's empirical reach.

### Figure 2

**Figure 2. The distillation model of authorship under the *mekiki* framework.**

Authorship operates as a repeating cycle in which the **Human Expert** evaluates the current state and sets the conditions for the next pass. **Evaluate** asks what is present, what is missing, and what should be adjusted. **Set Conditions** determines what to retain, what to discard, and which quality criteria apply. Together, these operations constitute specification, and human responsibility resides in them. The **Apparatus** (the AI system) reduces externalisation cost while processing **Raw Material** (collective knowledge, conventions, prior art, community practices, and data) in an environment that selectively attenuates interpersonal confounders. It produces an intermediate or final **Output**, which returns to the Human Expert for evaluation; every iteration is therefore human-mediated. The comparison below states the corresponding claim in text.

**What determines the output: the human, not the apparatus.**

|  | Different experts | Different AIs |
|---|---|---|
| Raw material | same | same |
| Conditions | different | same |
| Apparatus | same | different |
| **Output** | **different** | **same** |


---

