← All research

Foundation paper

From Strategic Intent to Verified Operating Change

A working design theory for the self-improving company

Executive proposition

This paper addresses an asymmetry that can arise in large-scale transformation. An organisation may be able to report spend, milestones, releases, licences, training and adoption while remaining less certain whether skilled capacity was genuinely recovered, whether work disappeared or moved elsewhere, whether AI reduced effort or converted execution into checking, or whether the new operating model is producing the intended result.

This is the problem Teho is being built to address.

The central thesis is that an organisation becomes progressively better at changing itself when it can maintain a continuing connection between five things: the outcome its leaders intend, the work that is actually being performed, the intervention it makes, the operating change that follows, and what that result should teach it about the next decision.

Many organisations already have sophisticated programme governance, financial reporting and continuous-improvement practices. The more specific claim here is that these mechanisms often remain fragmented. Strategy, operational evidence, implementation and benefits measurement may live in different systems, at different cadences and under different owners. A programme can therefore complete successfully while the organisation remains uncertain about what changed in the work, why it changed and whether the improvement will persist.

AI raises the importance of this problem. It allows organisations to alter the division of work between people and technology more quickly and at a much finer level. A new agent, data source, instruction, confidence threshold or review pattern can change how a team works without changing the formal process at all. The value does not reside in the model or agent alone. It depends on the resulting configuration of human judgement, automated execution, information, review, accountability and control.

The framework in this paper draws substantially on active inference. In its formal setting, active inference connects inference from partial observations, preferences over outcomes, policy selection and model revision [6,13]. This paper uses those concepts as a design framework for organisational change under uncertainty and adds a specifically organisational requirement: desired conditions, constraints and decision rights must be explicit and remain human-governed. That governance requirement is part of the Teho synthesis, not an empirical consequence of active-inference theory.

That does not mean a company is literally one biological or computational agent. For the purposes of this theory, the relevant unit is a bounded organisational subsystem that may comprise interacting human agents, AI agents where present, and other software, and that is steered through their interactions and the decisions of people with legitimate authority. Those participants have different information, capabilities, preferences and consequences. Strategic intent is governed, not discovered in telemetry. The unit of analysis is therefore a bounded team, workflow, function or value stream organised around an accountable human intent.

For an executive, the practical proposition is straightforward. Teho should help answer where human capacity and cost are going, which parts of work serve the function’s goals, where friction and lower-value activity consume time, how AI is actually being used, and where a change may create worthwhile capacity or performance improvement.

The canonical proposed loop is intent → observation → state inference → intervention → verification → learning. Its purpose is better-supported operating decisions, including a justified decision to continue, wait or retain a control. Retained experience supports that loop; an accumulating archive is not the proposition in itself.

“Verified” in the title means that a stated aspect of operating change has been checked against specified evidence within a defined scope. It is not a blanket certification that the intervention caused a business outcome or that its proposed mechanism is correct. Exposure, change in work, causal-effect confidence and mechanism support require distinct judgements.

Scope and status of the argument

This is a working design theory: it specifies a purpose, constructs, design principles and propositions for evaluation rather than reporting an empirically validated explanatory theory [20]. Consistent with action design research, evaluation should treat the product and its organisational setting as reciprocally shaping one another [21]. The design must therefore be tested in use, including how observation, interpretation and governance alter the work being studied.

Established literature provides foundations: enacted routines differ from formal representations [1]; organisations can be analysed as sociotechnical systems [8]; partially observable decision-process research distinguishes observation from inferred state [3]; causal claims require a distinction between observation and intervention [7]; and organisational learning research examines the creation, retention and use of experience [9–12]. The proposed connection between these ideas is Teho’s synthesis. The literature does not validate that synthesis, establish customer value or demonstrate that retained transition history improves later choices.

Teho is early-stage. Its current evidence foundation supports observation, human-assisted interpretation, recommendations and comparison over time; customers decide and implement changes. Section 7 separates that capability from the proposed connected architecture. The complete loop and its recurring value remain hypotheses, not completed empirical findings. The evaluation standard is whether the system improves consequential decisions and subsequent operating results with legitimate evidence and proportionate human effort.

1. The practical problem: transformation can remain open loop

Enterprise transformation normally contains a familiar sequence. Leaders define an outcome. Designers describe a future operating model. Programmes change technology, process, structure or capability. Operational systems report performance. Benefits and lessons are assessed later.

Each activity can be rigorous. The structural weakness is that they are not always maintained as one continuing account of change.

Strategic intent may sit in a business case. The current operating model is assembled through interviews, workshops, process maps and system data. The intervention becomes a programme plan or product backlog. Adoption is reported by the technology vendor. Operational results are measured elsewhere. Lessons are captured after the event in documents or in the memories of people who may move on.

In this failure mode, the organisation retains adjacent records rather than a connected model of intent, action and consequence.

That gap matters because delivery evidence is not the same as evidence of operating change. A programme may deliver scope while the intended work remains substantially unchanged. A tool may be adopted while its users recreate the old method around it. Automation may remove effort in one place and reappear as verification, exception handling or coordination elsewhere. Capacity may be declared released but disappear into lower-priority work rather than reach the customer, commercial or professional activity that justified the investment.

The distinction between open- and closed-loop control is useful here, provided it is treated as an analogy rather than a claim that a company is a deterministic machine. An open-loop system acts without using the resulting condition to choose its next action. A closed-loop system receives feedback, compares the observed condition with the desired condition and adjusts. Conant and Ashby’s good-regulator theorem is a conditional formal result, not evidence that an enterprise requires one central computational self-model; its relevant lesson here is the more modest one that effective intervention depends on an adequate representation of the system being regulated [2].

A programme asks: are we delivering the plan?

A continuing improvement system also asks: given the intended outcome, what appears to be happening in the work now; what transition did the intervention actually produce; what has that result taught us; and what should we do next?

The second question does not replace delivery discipline. It prevents delivery from becoming the sole proxy for improvement.

1.1 Why AI changes the management problem

AI changes work through more than a single implementation event. The allocation of labour between people and technology can move as models improve, instructions change, new tools and data become available, permissions expand, confidence thresholds shift and employees learn where a system can and cannot be trusted.

This makes a detailed target state more perishable. The strategic outcome may remain stable while the most sensible division of work changes during implementation.

AI can also create second-order effects. It may increase output faster than the surrounding operation can absorb it. It may transfer work from execution into checking. It may make normal cases faster while making exceptions harder to resolve. It may produce apparently released capacity that is subsequently consumed by new administration. These are changes to a sociotechnical system, not merely adoption of a tool.

AI may also lower the cost of the feedback loop itself. Models can help classify work, connect evidence, maintain hypotheses and prepare possible interventions. Whether they can do so accurately enough to improve important decisions is an empirical question, not an assumption.

The risk is therefore not only that organisations change too slowly. It is that they become able to change many local parts of the company without knowing whether the combined system is improving.

1.2 The target operating model becomes a trajectory

A target operating model remains useful when it expresses durable principles, responsibilities and constraints. It becomes brittle when it treats one detailed future configuration as fixed despite changing technology, economics and evidence.

The alternative proposed here is not permanent reorganisation. It is an adaptive trajectory with three horizons:

Stable intent: the outcome, values, constraints and accountable ownership that persist until deliberately revised.

Committed next transition: the bounded move the organisation is prepared to make now, supported by current evidence and explicit guardrails.

Provisional later path: possible subsequent moves, conditional on what the next transition reveals.

This resembles the receding horizon used in model-predictive control: plan beyond the immediate step, act within constraints, observe the result and recalculate [4]. The analogy supplies a useful discipline, not a promise that organisational behaviour is controllable in the engineering sense.

Continuous improvement, in this account, means continuous attention. The right next decision may be to scale, modify, constrain or stop an intervention; to preserve a valuable human practice; or to defer action until the evidence improves.

2. The representation problem: observation is not operating state

Organisational-routine research distinguishes an ostensive aspect—the abstract or generalised idea of a routine—from performative aspects—the specific actions undertaken by particular people at particular times [1].

Formal artefacts such as process maps, policies, system designs and role descriptions may express an intended routine, but they are not identical to its ostensive aspect. The practical distinction used in this paper is between those formal representations and the work actually enacted: the actions, judgements, workarounds, interruptions, handoffs, exceptions and local adaptations through which work is completed.

A transformation may change formal representations without producing a corresponding change in enacted work. For transformations intended to alter how work is performed, value depends on the enacted work changing in a useful way.

Yet enacted work is not directly available as one clean fact. Leaders receive partial observations: application activity, workflow events, agent traces, documents, outcomes, interviews and employee explanations. Each reveals something and omits something. Partially observable decision-process research offers one formal vocabulary for separating observation from hidden state, but it does not establish that an organisation is literally a POMDP [3].

Telemetry can show that a person moved between applications, repeatedly checked an output or spent time in a class of activity. It cannot by itself explain why the action mattered, whether it was avoidable, what knowledge it required, or whether the apparent inefficiency protected quality.

The appropriate output is therefore not a claim to complete organisational truth. It is a purpose-bound operating-state estimate: a bounded, evidence-backed and revisable account of the conditions relevant to one decision.

Such an estimate should be:

  • bounded to a defined team, workflow, function or value stream;
  • multidimensional rather than reduced to one productivity score;
  • traceable to the observations and explanations that support it;
  • explicit about coverage, confidence and missing evidence;
  • open to competing explanations and informed challenge;
  • sensitive to material variation between roles, cohorts and cases;
  • updated when new evidence or a new decision changes what matters.

Intent determines the scope and relevance of the estimate. It must not determine what the evidence is allowed to show. If the executive wants lower cost, that preference can specify the decision and constraints; it cannot make inconvenient work patterns disappear from the account.

2.1 Meaning is situated

The same observed activity can mean different things in different operating contexts.

Extensive human review may be avoidable burden when the objective is to reduce service cycle time. It may be an essential safeguard in a regulated remediation process. It may be a temporary learning mechanism while a team establishes where an AI agent is reliable. It may also indicate that poor source data is forcing reviewers to reconstruct context.

Observed activity does not contain its own managerial meaning.

For that reason, employee and operator explanation is not an optional courtesy added after the analysis. It is one source of evidence needed to interpret work. The system should expose hypotheses, not silently convert traces into judgement.

2.2 Privacy is part of the epistemology

Electronic-monitoring research reports context-dependent effects and, overall, a neutral association with performance alongside small adverse associations with strain and job attitudes [22]. It is therefore unsafe to assume that more granular observation is behaviourally or epistemically neutral.

Employees who perceive observation as punitive, individually evaluative or stripped of context may alter what they disclose or how they work. Whether, to what extent and under which conditions this degrades Teho’s evidence is an empirical question.

Teho’s proposition is that privacy-protective, team-level analysis can improve both legitimacy and evidential quality. That is a testable design hypothesis, not a settled result. It requires transparent purpose, proportionate collection, aggregation where possible, controlled access, human explanation and clear limits on use.

The aim is not to discover who worked hardest. It is to understand how a bounded system produces its outcome and where the configuration of people, agents, applications and controls may be improved.

3. The governance problem: a company is not one agent

Active inference offers a powerful language for the architecture proposed here. It also creates a seductive mistake: treating the company as though it had one mind, one objective and one clean boundary.

Companies do sense, interpret and act. They maintain models of customers, competitors and themselves. They establish goals, allocate attention, intervene, receive feedback and adapt. That resemblance makes it useful to describe an organisation as an adaptive system trying to move from its current condition towards preferred conditions under uncertainty.

Taken literally, however, the analogy becomes misleading.

A corporation is not one person scaled up. This paper models a bounded organisational subsystem—not the corporation as a whole—as a sociotechnical system that may include human agents, AI agents where present, and other software. It is steered through their interactions, within structures of authority, incentives and governance. Participants have different information, capabilities and local objectives; software constrains and enables possible action; and people remain accountable for legitimate purpose and consequential choices.

What appears to be one company-level action is therefore normally the result of hierarchy, negotiation, programmed behaviour, habit, local adaptation and incomplete coordination.

In active inference, an agent acts in relation to preferred states or outcomes. In a company, preference is not given by nature. It is governed.

The board may prioritise margin. A customer team may prioritise service quality. Risk may require stronger verification. Employees may value autonomy and manageable workload. Technology leaders may seek standardisation. A local manager may be rewarded for throughput even where enterprise value depends on end-to-end quality.

No software system can legitimately infer one corporate utility function and optimise it. The governing reference must be made explicit:

  • What outcome is being sought?
  • Which operating conditions would constitute a good result?
  • What constraints must remain true?
  • Who has authority to resolve the trade-offs?
  • Which affected groups can challenge the interpretation?
  • What evidence is strong enough for the decision at stake?

Human leadership does not disappear from the system. It becomes more clearly responsible for defining and revising intent.

3.1 The correct unit is a bounded multi-agent sociotechnical system

The useful unit is a bounded organisational system organised around an accountable intent.

It might be an account-management function seeking to increase time with customers; a remediation workflow trying to improve quality and throughput; a shared-service operation seeking lower cost without weaker control; or an onboarding process trying to increase speed while preserving due diligence.

The boundary is practical rather than metaphysical. It identifies the people, any participating AI agents, and the other software shaping their possible actions; the work and outcomes in scope; the available evidence; the interventions the accountable owner can authorise; and the adjacent systems where consequences may appear.

Even within that boundary, the model should preserve material local variation. A team can show high mean AI adoption while one role carries all the correction work. A workflow can improve in normal cases and fail at the exceptions that matter most. An intervention can improve an executive metric by transferring effort, uncertainty or risk to a less visible group.

Aggregation can protect privacy and support system-level thinking. It should not average away consequences that change the decision.

An AI agent does not merely replace a task. It alters the distribution of information, execution, judgement, review, authority, accountability, skill, exception handling and trust. Software systems also shape the environment within which both human and AI agents perceive and act. The unit of improvement is the resulting configuration, not the isolated model, person or task.

3.2 Resistance can be information

Organisational-change research has challenged the treatment of resistance as a unitary barrier: responses to change may be ambivalent, and interactions between change agents and recipients can reveal defects in the intervention or its implementation [23,24]. A systems view therefore asks what a resistant response may reveal before classifying it simply as poor adoption.

Employees may be protecting customer quality. A workflow may depend on information the redesign removed. Incentives may still reward the old behaviour. A downstream team may reject a new output because it cannot trust its provenance. The technology may require more cognitive effort than the task it replaces.

Not every objection is correct, and no organisation can avoid difficult decisions. But a system that returns towards its prior pattern may be revealing a stabilising dependency, incentive, routine or constraint. Diagnosing that evidence is more useful than treating all deviation as poor adoption.

4. Active inference as a design framework

In the formal literature, active inference connects inference about hidden states, preferences over outcomes, policy selection and learning within a generative model [6,13]. The discrete-state synthesis offers a general framework in which the assumptions of particular models can be made explicit, not a replacement for all other models [6]. Its use here is similarly specific: separate what observations suggest about current conditions from what competing actions are expected to change, then consider both preferred consequences and the information those actions could produce.

Applied carefully to a bounded organisational system, the mapping is:

Observations: authorised evidence of enacted work, agent behaviour, workflow events, operator explanation and relevant outcomes.

Hidden operating conditions: the purpose-bound state of the work that leaders cannot observe directly—for example, where skilled capacity is being consumed, where review burden sits, or why a new tool is not changing the workflow.

Observation model: an account of how the available evidence could arise from underlying operating conditions. Repeated checking, for example, could reflect poor source access, weak output quality, a necessary control or a temporary learning practice. The same visible pattern can therefore support several explanations.

Action-dependent transition model: an account of how those conditions might change under each candidate intervention, including expected delays, side effects and uncertainty. Source access, training and a review-policy change make different predictions about what should happen next. This model is distinct from the observation model; together they supply parts of the proposed generative account.

Preferred conditions: the outcome, operating characteristics and constraints defined through human governance.

Policies or actions: candidate organisational interventions, not autonomous commands.

Pragmatic value: the expected contribution of an intervention to the desired operating and business outcomes.

Epistemic value: the information an intervention is expected to produce about the state or transition mechanism.

Prediction–outcome discrepancy: evidence that can cause people and the system to revise the model used for the next decision.

The practical design choice is to retain competing explanations and their different predictions until evidence warrants narrowing them. Candidate interventions are compared not only by expected immediate improvement but by whether their possible outcomes would discriminate between those explanations. Information has value when it can improve a consequential later choice; collecting more data is not inherently useful.

This does not imply that contested organisational values can be collapsed into one mathematical objective. Cost, risk, employee impact, quality and control may not be commensurable. Accountable humans must set constraints, retain vetoes and judge trade-offs. Formal notation can clarify the architecture; it cannot confer legitimacy on a decision.

Nor does the framework establish that a company possesses one unitary generative model or a biologically meaningful Markov blanket. The group- and organisation-level sources cited here are conceptual or based on simplified simulations [14–16], while critiques of broad free-energy and Markov-blanket claims show why additional assumptions matter [17,18]. The analogy earns its place only if it improves the representation of uncertainty, the design of interventions and the quality of later decisions.

The substantive Teho proposition is that a bounded, human-governed organisational system can learn more deliberately when incomplete evidence of enacted work is connected to explicit preferred conditions, diagnostic interventions and interpreted prediction–outcome discrepancies. The following example shows the proposed design in use.

4.1 A worked diagnostic decision

This example is wholly hypothetical. It illustrates a decision protocol, not a Teho customer, product result or estimated effect.

A customer-service team has adopted an AI assistant for recurring enquiries, but the manager sees little usable capacity returning to customer work. Staff use the assistant and still spend substantial time checking and reconstructing answers. Quality must be maintained, existing access restrictions must remain in force, and any recovered capacity should reduce the queue rather than become additional administration.

Two explanations remain plausible. The skill-and-prompt explanation is that staff cannot reliably elicit useful answers, so drafting creates avoidable correction. The source-access-and-workflow explanation is that the assistant lacks reliable access to an approved source: staff reconstruct context and repeat checks regardless of prompt quality. A further possibility is that the existing review process remains necessary even when answers improve. The observed checking pattern alone does not select among these accounts. Staff explanation and case review are needed to establish what is being checked and why.

The accountable manager compares three interventions.

  • Broad prompt training could improve output if skill is the main constraint. It requires staff time and may be useful, but a mixed result would be hard to interpret: attendance, task mix and use of the new method may vary, while source access remains unchanged. Its diagnostic value for the present uncertainty is limited.
  • A bounded source-access intervention would give the assistant access to one already approved knowledge source for an eligible, recurring enquiry type, within existing user access rights. Prompts, training, model version and review rules would remain fixed. Its immediate benefit is narrower, but its predictions are clearer: if missing context drives reconstruction, that work should fall among genuinely exposed cases. It is reversible by withdrawing the added access, subject to the existing access and security review.
  • Removing a review step might release capacity quickly if review is redundant. It would also expose the team to quality risk and could conceal rather than resolve the source of checking. Because the evidence does not yet justify that risk, the manager does not select it.

The manager chooses the bounded source-access intervention for its combination of possible benefit, diagnostic value and reversibility. Eligible cases are assigned to the existing or source-enabled configuration using a prespecified random allocation where operationally feasible. Case eligibility, assignment, outcome measures, review period and stop conditions are fixed before the test. If allocation cannot be maintained, the comparison is treated as less informative, not quietly described as an experiment. The team records concurrent changes and checks for learning or workarounds spilling between configurations.

Exposure is checked before outcomes are interpreted. The source must have been available with the intended permissions, retrieved in the relevant cases and used in the workflow—not merely enabled in a release. Authorised system evidence and sampled case review establish that account without turning team analysis into individual productivity ranking. Missing exposure records or departures from the agreed configuration limit what can be concluded.

Suppose, hypothetically, exposed cases show less manual source reconstruction and checking work than comparable cases in the existing configuration, while the prespecified quality checks show no deterioration over the review period. That pattern weakens a skill-only explanation and supports the view that access to context matters. It does not show that skill is irrelevant, that all checking is unnecessary or that the result will persist at scale. Assignment integrity, spillovers, case mix and the sensitivity of the quality checks still need scrutiny. Even a credible effect of the access intervention would not by itself establish the exact mechanism; case evidence is needed to distinguish reduced reconstruction from other changes in how staff use the assistant.

The next decision is therefore to extend the same bounded configuration to a further comparable enquiry type while retaining the review control, not to launch broad retraining or remove review across the team. The revised account says that source access is a supported contributor in the tested setting; it also preserves the remaining uncertainty about generalisation and the review requirement.

Other outcomes would change that decision. No meaningful exposure would call for an implementation repair and retest, not rejection of the hypothesis. Verified exposure without reduced reconstruction would weaken the source-access explanation and make a separately scoped skill or workflow investigation more useful. Less reconstruction without usable capacity returning would direct attention to where the time went. Worse quality would trigger the agreed stop or rollback regardless of any time saving. An inconclusive comparison could justify waiting for adequate evidence rather than selecting a winner.

Good conventional experimentation or continuous-improvement practice could reach the same choice. The active-inference-informed contribution is the explicit connection between competing state explanations, action-dependent predictions, practical and informational value, and the next model update. Whether requiring that connection improves decisions enough to justify its cost is an evaluation question. The example demonstrates the design choice; it does not establish theoretical exclusivity or product effectiveness.

5. The intervention-learning problem: change as a hypothesis

An intervention is a claim about transition.

If an organisation introduces an AI assistant, changes a workflow, removes an approval or reorganises a team, it is implicitly claiming that the current conditions will respond in a particular way.

A useful intervention record should make that claim explicit:

  • the operating gap being addressed;
  • the evidence and competing explanations;
  • the proposed action;
  • the expected mechanism;
  • the predicted change in work;
  • the desired business consequence;
  • the people or cases expected to be exposed;
  • the owner, approvals and decision rights;
  • the guardrails and stop conditions;
  • the expected time to signal;
  • the cost, risk and reversibility.

This does not turn leadership into laboratory science. It prevents assumptions from disappearing once the work begins.

5.1 Improvement value and learning value

An organisation often has to act before it understands the transition mechanism with confidence. One action may offer the largest expected benefit if the current hypothesis is correct. Another may produce a smaller direct benefit but resolve uncertainty that affects many later choices.

Dual control and active inference both give this trade-off a formal home [5,6]. In practice it should be treated as a human-governed, multi-criteria decision rather than a literal utility calculation.

The decision should consider expected progress, information gain, implementation cost, operational and employee risk, reversibility and time to signal. When uncertainty and downside are high, a bounded and reversible intervention is often preferable. Where the mechanism is well established and delay is costly, larger action may be justified.

The principle is not “make every change small”. It is “make the scale and irreversibility of the change proportionate to the evidence and downside”.

5.2 Actual exposure comes before outcome

A change cannot be evaluated from its launch date.

The organisation first needs to establish what was actually put into practice: who or what was exposed; how often the new workflow or agent was used; which variant was present; whether the intended process was followed; and what concurrent changes occurred. Evidence that the intervention reached the operation is not yet evidence that its proposed mechanism operated.

No outcome movement may mean that the intervention was ineffective. It may also mean that the intervention never meaningfully reached the operation. Those are different conclusions.

5.3 Verification without causal overreach

Three evidential judgements must remain separate, alongside the descriptive question of what changed:

  1. Exposure verification: what intervention actually reached which people or cases, in what form, and with what fidelity to the intended configuration?
  2. Causal-effect confidence: how credible is the claim that the intervention made a difference to the observed work or outcome, relative to what would otherwise have happened?
  3. Mechanism support: what evidence supports the explanation of how that difference arose, and which rival explanations remain plausible?

A before-and-after comparison can describe a difference; it does not by itself identify an intervention effect. Timing, workload, case mix, other changes and measurement drift may account for the movement. Causal inference’s distinction between observation and intervention makes the need for an explicit counterfactual and its assumptions clear [7].

Evidence can range from temporal association to comparative designs and credible experimental or quasi-experimental attribution. A matched comparison, staggered rollout or interrupted time series is not stronger merely because it has a methodological label. Its value depends on assignment, baseline comparability, trends, spillovers, co-interventions and other relevant identification assumptions. Where those assumptions cannot be defended, the claim must remain correspondingly limited.

Mechanism support is a separate dimension, not another rung on the same ladder. Process evidence, timing and operator explanation may support a plausible pathway without ruling out alternative causes of the outcome. Conversely, a well-designed experiment may identify an effect of an intervention package while leaving the pathway uncertain. Changing training, source access and review rules together might test the package; it would not identify which component produced the result. Reusing that package as though one mechanism had been established would carry unsupported certainty into the next decision.

The record should therefore state the observed change, exposure evidence, effect claim and its assumptions, and mechanism evidence separately. Unintended consequences and unresolved alternatives belong in the same account. The required evidence should be proportionate to the stakes, and the language should identify what has and has not been established.

The difference between predicted and observed consequences is not organisational learning by itself. It becomes learning when it is interpreted, challenged, retained and used to support a future decision—including a warranted decision to keep the present course.

6. Transition Memory: making change cumulative

The design problem addressed here is not that organisations lack records of change. They may preserve business cases, operating models, process designs, training materials, issue logs and benefits reports. The proposed gap is narrower: those records may not preserve, in a reusable form, the relationship between a prior state, an intervention hypothesis, actual exposure, observed consequence, confidence and the implication for the next decision.

Organisational-memory and organisational-learning research already examines how experience is acquired, retained and used [9–12]. This paper calls the narrower intervention-linked construct proposed here Transition Memory.

A transition record contains:

Before: intent; scope; prior operating-state estimate; evidence and uncertainty; competing hypotheses; context and baseline.

Change: intervention; expected mechanism and competing pathways; predicted transition; owner and approvals; intended population; actual exposure and its verification; guardrails; cost, risk and reversibility; concurrent changes.

After: observed transition; operational and business outcomes; unintended effects; causal-effect confidence and identification assumptions; mechanism support and unresolved alternatives; discrepancy between prediction and result; interpretation; model update; next decision.

Repeated records could allow an organisation to ask:

  • Which interventions worked in comparable conditions?
  • What conditions appeared necessary?
  • Where did burden move?
  • Which assumptions repeatedly failed?
  • How long did meaningful signals take to appear?
  • Which roles or cohorts responded differently?
  • Why did a successful pilot fail to scale?

This is not a claim that experience automatically compounds. Learning requires that the original hypothesis was explicit, the intervention and exposure were recorded accurately, uncertainty was preserved, the result was interpreted, and the resulting knowledge was retrievable and relevant to a later decision. Retrieval must preserve the distinction between a verified deployment, evidence of an intervention effect and support for its mechanism. Otherwise a plausible explanation can harden into a rule merely through repetition.

The strongest early learning is likely to be company-specific. Systems, culture, decision rights, controls and workforce composition condition what happens. A transition that succeeded elsewhere is not automatically transferable.

Higher-level patterns may eventually be reusable across customers—capture methods, semantic descriptions, common mechanisms, diagnostic interventions and calibrated evidence patterns—but only within privacy, contractual and governance boundaries.

Transition Memory is therefore a supporting learning mechanism whose incremental value must be tested. Its usefulness depends on improving later choices, not on the volume of history retained.

7. What Teho is building

The architecture connects an existing evidence foundation to a proposed decision-and-learning circuit. The distinction matters for evaluation: a useful baseline does not establish the value of the complete system, and the complete system does not establish the incremental value of retained history.

7.1 The evidence foundation

Teho can establish a passive baseline from work visible on approved employee devices, together with supported AI-agent session or event data. It is designed to protect individual privacy and can begin without a mandatory application-integration project.

The purpose is not to monitor individual productivity. It is to make the collective shape of work legible: where team time and cost are concentrated; how capacity is divided between judgement, mechanics, administration, coordination and rework; where friction and checking burden sit; how AI is being used; and which activities may warrant protection, redesign, delegation or automation.

The baseline is an observation and diagnosis capability. It can challenge a formal process, a transformation assumption or an executive belief about where capacity is going. Human context remains necessary because observed activity does not explain its own purpose.

In terms of the canonical loop, this supports observation and human-assisted state inference, preparation of recommendations and comparison over time. The output is a bounded, evidence-backed account of work. Its value depends on whether it improves the evidence available for a real operating decision.

7.2 The smallest complete circuit

The next product layer would connect the baseline explicitly to a live operating decision.

An accountable executive defines the desired outcome, the system in scope, the conditions that would constitute improvement and the constraints that must remain true. Current evidence is interpreted in relation to that intent.

A chosen change is preserved as a governed transition hypothesis: given what appears to be happening now, this intervention is expected to alter these aspects of work, through this mechanism, within these guardrails and over this period.

The proposed circuit would then establish whether the change was actually put into practice, connect subsequent work and outcome evidence to it, record unintended effects, distinguish effect confidence from mechanism support, and state what the result warrants for the next recommendation.

Teho does not need to execute the intervention itself to prove this circuit. The immediate requirement is to connect intent, prior state, chosen intervention, actual exposure, subsequent state and learning in one trustworthy record.

7.3 The connected architecture

The proposed architecture is a living, cross-vendor account of how a bounded organisational system operates and changes. Its six functions follow the canonical loop:

  1. Intent: outcome, preferred conditions, constraints, scope, horizon and decision rights.
  2. Observation: authorised, privacy-protective evidence of work across people, supported agents, applications and relevant outcomes.
  3. State inference: a bounded and contestable estimate of what is happening now, with confidence and competing explanations.
  4. Intervention: comparison and human selection of candidate transitions by expected improvement, learning value, cost, risk, reversibility and time to signal.
  5. Verification: evidence of actual exposure and consequence following customer-owned implementation, with separate judgements about causal effects and mechanisms. Any later execution of bounded approved actions through existing systems would require explicit authority, guardrails and rollback.
  6. Learning: retention of the transition record and revision of the state and action-dependent transition models used for the next decision.

The destination is not autonomous company-wide management. It is a human-governed system that makes operating change more observable, explicit, testable and cumulative.

7.4 The proposition under test

The proposition is that these functions can be maintained as one useful product rather than reconstructed repeatedly through consulting.

That requires more than collecting data. The system must produce interpretations that operators recognise, better-supported decisions, transition records that survive scrutiny, and learning that improves a later recommendation or its justification.

Human explanation, correction and governance remain part of the architecture because the system is modelling purposeful work, not only machine events.

8. Where the customer value comes from

The proposed value does not come merely from creating more organisational data. It comes from reducing uncertainty around consequential operating decisions.

8.1 A clearer account of capacity and cost

An executive may know the total cost of a function while lacking a defensible activity-level view of what that cost is buying.

Teho can help show how skilled time is distributed: how much reaches the high-value work associated with the function’s goals; where it is consumed by mechanics, coordination, checking or rework; and where friction or fragmented information absorbs capacity.

This creates an identifiable value pool. It does not mean every lower-value activity can or should be removed. It enables the accountable leader to price the current work and decide which activity should be protected, redesigned, delegated, automated or stopped.

8.2 Better diagnosis before further investment

A company may interpret low AI adoption as a training problem when the constraint is missing data, weak output quality, poor workflow design, inadequate permissions or the fact that using the tool takes longer than completing the task directly.

Acting on the wrong explanation can consume investment while leaving the work unchanged. A purpose-bound state estimate cannot eliminate uncertainty, but it can expose competing explanations and support a more diagnostic next move.

The economic value may come from avoiding a broad, badly targeted programme as much as from identifying a new intervention.

8.3 More credible evidence after change

A transformation benefit may be claimed through adoption, estimated time saved or an outcome that improved after implementation. None, alone, establishes that the intended work transition occurred or that the intervention caused the result.

By preserving the prior state, expected mechanism, actual exposure and observed consequences, Teho could help distinguish a change that worked, a change that was never meaningfully enacted and a change whose apparent benefit has plausible alternative explanations.

This does not guarantee causal proof. It improves the integrity of the decision to scale, alter or stop.

8.4 Capacity that reaches a valuable destination

Removing work does not automatically create value. Capacity can be absorbed by more review, new administration or lower-priority activity.

The relevant question is not only whether time was saved, but what the operation became: whether skilled capacity moved towards customer, commercial, professional or control activity that the accountable executive values, without moving unacceptable cost or risk elsewhere.

8.5 Better later decisions

A baseline can be valuable as a one-off diagnosis. A continuing system becomes more valuable if it retains what happened when the organisation acted.

The customer would not continue using Teho merely because a dashboard needs refreshing. It would continue because the operating state changes, new interventions are made and the accumulated transition history makes later decisions better informed.

The economic case must ultimately be demonstrated in customer terms: capacity recovered; avoidable cost removed; higher-value activity increased; transformation spend redirected; risk better controlled; or an important decision made with greater confidence. Teho should not claim these outcomes simply because it can observe work. It must show the chain from evidence to decision, decision to intervention and intervention to credible operating consequence.

9. The Teho bet

Teho’s bet is a chain of propositions. Every link matters.

First, AI will make the operating model more dynamic. As models, agents and workflow tools change, organisations will repeatedly reconsider what people should perform, what software should execute and where human judgement, review and accountability should remain. The management problem moves from implementing one stable future-state design towards steering a continuing reconfiguration of work.

This emphasis on repeated sensing, action and reconfiguration is adjacent to dynamic-capabilities research, but that literature does not establish the Teho architecture or its product value [19].

Second, the information needed to understand that reconfiguration can remain fragmented across systems and owners. Financial systems show outcomes. Workflow platforms show formal process. Application and agent logs show events. Surveys and interviews provide interpretation. The proposed gap is the continuing connection between strategic intent, enacted work, intervention and consequence—not the absence of useful existing methods.

Third, passive, privacy-protective observation may make enough of the hidden operating condition visible to improve a real executive decision. The standard is not a complete digital twin. It is a bounded and contestable account that is more useful than the information previously available.

Fourth, active-inference principles may improve the product architecture. Partial observations support explicitly qualified beliefs about current conditions. Human-governed intent defines preferred conditions and constraints. Candidate actions are compared through their different transition predictions, expected improvement and information value. Differences between prediction and result prompt revision of the relevant model. Whether the resulting judgements are well calibrated must itself be evaluated.

If this framework produces no better decisions than conventional analytics, continuous-improvement practice or experienced consulting, the theoretical synthesis has not created product value.

Fifth, Transition Memory may become a durable organisational asset. Individual models, agent platforms and execution tools will change. A company’s structured knowledge of what happened when it changed its own work may persist across those technologies.

That knowledge may contain conditions a generic model cannot infer: where delegation succeeds; which controls create useful oversight; where automation shifts effort into verification; and which local routines are essential to quality.

No single part constitutes a defensible advantage. Potential defensibility would come from the combination: evidence of work as enacted across different software; human and supported-agent activity interpreted in one model; privacy controls applied close to collection; analysis conditioned by accountable intent; visible provenance and confidence; and accumulated company-specific transition history.

Sixth, Teho must turn this design into a repeatable and recurring product. If every customer requires the founder or an analyst to reconstruct the meaning manually, Teho may remain a consultancy with useful software. If the product can increasingly produce credible state estimates, transition records and relevant learning with proportionate human correction, the economics and scalability change.

The commercial and investment hypothesis is that a repeatable connection between these functions can create recurring customer value. It depends on evidence at each layer: observation, state validity, decision value, transition evidence, learning and recurrence. Retained history matters only insofar as it strengthens that operating capability.

10. Evaluation and failure conditions

A working thesis should identify results that weaken it rather than reinterpret every outcome as support.

10.1 A prospective evaluation approach

Evaluation should begin with bounded, consequential decisions for which there is an accountable owner, an observable population and a plausible period in which consequences can emerge. Before presenting Teho’s account, record the decision that would otherwise be made, its evidence, assumptions, alternatives and confidence. The comparator should be credible current practice: the organisation’s existing analytics, continuous-improvement process or experienced advisory support, with its available information and effort recorded. Comparing against an artificially uninformed manager would not test the proposition.

First test the connected system against that comparator, not the memory feature in isolation. Where feasible, assign comparable decision episodes or operating units prospectively to the two approaches, with common constraints and outcome definitions. Where random allocation is impractical, use an explicit comparison design and preserve its limitations. The design must account for selection, differences in task difficulty and knowledge spilling between groups; otherwise apparent decision value may reflect where or by whom the system was used.

Before each evaluation begins, specify the decision horizon, follow-up horizon, success and harm measures, evidence requirements, relevant costs and conditions under which the result will be inconclusive. The horizon must fit the work: an immediate workflow signal and a durable business consequence may need different observation periods. It should not be extended after an unfavourable result merely to find a favourable one. These are proposed studies, not completed tests; thresholds should be agreed for the decision context rather than invented here.

The major propositions require related but distinct tests:

  • Observation and interpretation: compare the accuracy, coverage and decision relevance of the operating account with current evidence, using sampled source checks and structured operator challenge. Record disagreements, missing populations, privacy constraints and the staff or analyst time needed to reach an acceptable account. Accurate traces with no additional decision value, unresolvable misinterpretation or unacceptable collection requirements would weaken this proposition.
  • Decision support and diagnostic value: assess whether the recommendation is better supported, not merely different. Use prespecified criteria for evidential support, explicit alternatives, constraint compliance, treatment of uncertainty and whether the proposed action can answer the question it is meant to test. Independent reviewers, blinded to the source of a recommendation where practical, can assess its justification before outcomes are known. Follow up whether the expected information was actually obtained and used. Count a well-supported continuation, wait or retained control as a valid decision; count unnecessary intervention, unjustified confidence and avoidable decision delay as failures.
  • Operating consequence: verify exposure first, then compare relevant work, quality, risk and capacity-destination outcomes over the agreed follow-up period. Track burden transferred to other teams and the costs of implementation, interpretation and verification. Report observed change, causal-effect confidence and mechanism support separately. A persuasive narrative without usable exposure evidence, a benefit erased by downstream burden or a claim stronger than the design permits would not pass.
  • Retained-history contribution: only after testing the connected system, compare later recommendations made with and without access to relevant prior transition records, holding current evidence, decision task and available effort as similar as practicable. Independent assessment can test whether history improves the justification, avoids a known error or appropriately changes confidence. Prospective follow-up is still needed for operating consequences. Retrieval that adds obsolete assumptions, false analogies, additional cost or confidence without support would count against the memory hypothesis.
  • Trust, applicability and recurrence: across successive decisions, record whether legitimate access and participation can be sustained, which workflow types support useful inference, and how much human correction, explanation and support remain necessary. Measure recurring decision use and net value rather than dashboard visits. A valuable first diagnosis followed by no worthwhile second use, or costs that remain disproportionate to the decisions supported, would limit the product and commercial claims.

Decision quality and outcome quality should not be collapsed. A reasonable decision can have an adverse result under uncertainty; a poorly justified decision can succeed by luck. Prospective assessment of the rationale, followed by outcome and cost evidence, reduces the temptation to let either a persuasive explanation or a favourable result validate the entire system. Changes to the product during field evaluation should be versioned and reported, consistent with a design being refined in its organisational setting [21].

10.2 Theoretical failure

Active inference, control theory and partial-observability language may add terminology without adding explanatory or decision value. If the framework merely renames an ordinary improvement cycle, the synthesis is not distinctive. It earns its place only if it improves how hidden state, uncertainty, preferred conditions, information-seeking action and prediction–outcome discrepancy are represented and used.

10.3 Observational failure

Teho may capture accurate patterns of activity without adding material evidence to what executives, operators or employees already know. A confirmed decision can still gain value from stronger support or the removal of a consequential uncertainty. If the account improves neither the decision nor its justification enough to warrant its cost, observation alone is not sufficient product value.

10.4 State-inference failure

The product may still depend on extensive bespoke analyst interpretation, or operators may reject the resulting account as a poor representation of their work. A state estimate is useful only if it is bounded, recognisable, evidence-backed and open to correction.

10.5 Decision failure

A clear account of current conditions may be interesting without supporting a better choice about what to prioritise, protect, stop, redesign, test or leave unchanged. A changed decision is not automatically an improvement. If Teho adds intervention, delay or confidence without better justification than the comparator, it has not closed the gap between insight and decision value.

10.6 Transition failure

The organisation may make a change, but Teho may be unable to establish actual exposure or what relevant aspects of work changed. Mechanism evidence may also remain inadequate. If the record cannot distinguish these gaps from evidence of an ineffective intervention, it cannot reliably support the next decision.

10.7 Causal failure

Telemetry may encourage confident claims the evidence cannot support. If Teho converts temporal association into declared return on investment, trust will deteriorate. Alternative explanations must remain visible and the strength of the language must match the strength of the evidence.

10.8 Memory failure

The organisation may accumulate transition records that do not improve a later recommendation or its evidential support. If prior transitions are not comparable, retrievable or sufficiently informative to justify the cost of using them, Transition Memory becomes another archive. A changed recommendation based on a false analogy is worse than no reuse.

10.9 Trust-and-utility failure

Team-level evidence may prove too coarse for useful inference, while more granular observation may require forms of individual visibility that customers and employees will not accept. The thesis depends on demonstrating that legitimate observation can still produce decision-grade evidence.

10.10 Applicability failure

The framework may work in bounded, recurring workflows with observable populations and relatively prompt outcomes but fail in highly creative, unique or long-horizon work. That would not make a narrower product valueless, but it would limit claims about a company-wide operating system.

10.11 Commercial failure

Customers may value an initial baseline but not make a second consequential decision with Teho. Delivery may remain too bespoke. Correction and support costs may not fall. Transition history may not create enough additional value to support an annual relationship.

These tests are not peripheral. They define the route by which the thesis can become credible. Theoretical coherence may justify the experiment; only observed customer decisions, transitions and recurrence can justify the company built around it.

11. Counterarguments and limitations

The theory should be constrained by several objections.

Organisations are not controllable machines. They are reflexive, political and path-dependent. A model can influence the people it describes. Control-theory language must remain a design discipline, not a claim of deterministic steering.

There is no complete organisational state. Any state estimate is purpose-bound. What is relevant to one decision may omit what matters to another.

Not all valuable work is observable. Tacit knowledge, relationship quality, judgement and informal coordination can resist measurement. Human explanation remains necessary.

Causal attribution may remain weak. Concurrent changes, small samples and long lags may limit the strength of conclusions.

Learning may not transfer. A pattern that holds for one context may fail in another. Company-specific memory is the more defensible first claim.

Privacy constrains granularity. This is a legitimate limit, not an inconvenience to be engineered away.

Existing methods already close parts of the loop. Continuous improvement, process and task mining, transformation governance, digital adoption and organisational research each address parts of the problem. Teho’s claim is not that they do nothing. It is that the evidence and learning can remain fragmented, and that enacted work across people and agents could be connected more directly to human-governed intent and repeated intervention.

The value may be bounded. The approach may be most effective where there is a definable accountable outcome, a recurring pattern of work, an observable population and a consequential next decision.

These limitations do not invalidate the design. They define where it should be tested and where its claims must stop.

12. Conclusion

The self-improving company is not an autonomous company.

It is an organisation that becomes better at changing itself under human intent and authority. It can maintain a useful account of enacted work; make the assumptions behind an intervention explicit; observe whether the change was actually put into practice; distinguish changed conditions from confident causal stories; and retain what the result should change about the next decision.

The active-inference-informed design connects two questions that can otherwise drift apart: what might explain the work we observe, and what would each possible intervention change or teach us? Its practical contribution is to make those explanations and action-dependent predictions explicit, compare improvement with diagnostic value, and use the result to revise the next decision. Human governance supplies the legitimate preferences, constraints and authority. Conventional methods can share these disciplines; the proposed synthesis must earn its place through better-supported decisions and worthwhile operating consequences.

This paper models the relevant organisational subsystem—not the company as a unitary agent—as a multi-agent sociotechnical system comprising people, any AI agents in use, and other software. Meaning, legitimate purpose and accountability remain human. The model is bounded and contestable. Privacy is a design constraint and a proposed source of better evidence. Causal confidence is graded rather than assumed.

The essential test is a complete, trusted circuit: an accountable intent, an interpretable account of work, a justified choice, evidence of what was enacted and what followed, and a better-supported next decision. Retained transition knowledge supports the repetition of that circuit. It is a means by which improvement may become cumulative, not a substitute for showing that the work improved.

That bet is worth testing because the alternative is increasingly uncomfortable: organisations acquiring ever more powerful means of changing work while remaining unable to explain, with confidence, what those changes actually did.

References

[1] Feldman, M. S., and Pentland, B. T. (2003). “Reconceptualizing Organizational Routines as a Source of Flexibility and Change.” Administrative Science Quarterly, 48(1), 94–118. https://doi.org/10.2307/3556620

[2] Conant, R. C., and Ashby, W. R. (1970). “Every Good Regulator of a System Must Be a Model of That System.” International Journal of Systems Science, 1(2), 89–97. https://doi.org/10.1080/00207727008920220

[3] Kaelbling, L. P., Littman, M. L., and Cassandra, A. R. (1998). “Planning and Acting in Partially Observable Stochastic Domains.” Artificial Intelligence, 101(1–2), 99–134. https://doi.org/10.1016/S0004-3702(98)00023-X

[4] Mayne, D. Q., Rawlings, J. B., Rao, C. V., and Scokaert, P. O. M. (2000). “Constrained Model Predictive Control: Stability and Optimality.” Automatica, 36(6), 789–814. https://doi.org/10.1016/S0005-1098(99)00214-9

[5] Filatov, N. M., and Unbehauen, H. (2000). “Survey of Adaptive Dual Control Methods.” IEE Proceedings — Control Theory and Applications, 147(1), 118–128. https://doi.org/10.1049/ip-cta:20000107

[6] Da Costa, L., Parr, T., Sajid, N., Veselic, S., Neacsu, V., and Friston, K. (2020). “Active Inference on Discrete State-Spaces: A Synthesis.” Journal of Mathematical Psychology, 99, 102447. https://doi.org/10.1016/j.jmp.2020.102447

[7] Pearl, J. (1995). “Causal Diagrams for Empirical Research.” Biometrika, 82(4), 669–688. https://doi.org/10.1093/biomet/82.4.669

[8] Trist, E. L., and Bamforth, K. W. (1951). “Some Social and Psychological Consequences of the Longwall Method of Coal-Getting.” Human Relations, 4(1), 3–38. https://doi.org/10.1177/001872675100400101

[9] Argote, L., Lee, S., and Park, J. (2021). “Organizational Learning Processes and Outcomes: Major Findings and Future Research Directions.” Management Science, 67(9), 5399–5429. https://doi.org/10.1287/mnsc.2020.3693

[10] March, J. G. (1991). “Exploration and Exploitation in Organizational Learning.” Organization Science, 2(1), 71–87. https://doi.org/10.1287/orsc.2.1.71

[11] Walsh, J. P., and Ungson, G. R. (1991). “Organizational Memory.” Academy of Management Review, 16(1), 57–91. https://doi.org/10.5465/AMR.1991.4278992

[12] Zollo, M., and Winter, S. G. (2002). “Deliberate Learning and the Evolution of Dynamic Capabilities.” Organization Science, 13(3), 339–351. https://doi.org/10.1287/orsc.13.3.339.2780

[13] Friston, K. (2010). “The Free-Energy Principle: A Unified Brain Theory?” Nature Reviews Neuroscience, 11, 127–138. https://doi.org/10.1038/nrn2787

[14] Kaufmann, R., Gupta, P., and Taylor, J. (2021). “An Active Inference Model of Collective Intelligence.” Entropy, 23(7), 830. https://doi.org/10.3390/e23070830

[15] Waade, P. T., et al. (2025). “As One and Many: Relating Individual and Emergent Group-Level Generative Models in Active Inference.” Entropy, 27(2), 143. https://doi.org/10.3390/e27020143

[16] Fox, S. (2021). “Active Inference: Applicability to Different Types of Social Organization Explained through Reference to Industrial Engineering and Quality Management.” Entropy, 23(2), 198. https://doi.org/10.3390/e23020198

[17] Biehl, M., Pollock, F. A., and Kanai, R. (2020). “A Technical Critique of Some Parts of the Free Energy Principle.” arXiv:2001.06408. https://doi.org/10.48550/arXiv.2001.06408

[18] Bruineberg, J., Dolega, K., Dewhurst, J., and Baltieri, M. (2022). “The Emperor’s New Markov Blankets.” Behavioral and Brain Sciences, 45, e183. https://doi.org/10.1017/S0140525X21002351

[19] Teece, D. J. (2007). “Explicating Dynamic Capabilities: The Nature and Microfoundations of (Sustainable) Enterprise Performance.” Strategic Management Journal, 28(13), 1319–1350. https://doi.org/10.1002/smj.640

[20] Gregor, S., and Jones, D. (2007). “The Anatomy of a Design Theory.” Journal of the Association for Information Systems, 8(5), 312–335. https://doi.org/10.17705/1jais.00129

[21] Sein, M. K., Henfridsson, O., Purao, S., Rossi, M., and Lindgren, R. (2011). “Action Design Research.” MIS Quarterly, 35(1), 37–56. https://doi.org/10.2307/23043488

[22] König, C. J. (2025). “Electronic Monitoring at Work.” Annual Review of Organizational Psychology and Organizational Behavior, 12, 321–342. https://doi.org/10.1146/annurev-orgpsych-110622-060758

[23] Piderit, S. K. (2000). “Rethinking Resistance and Recognizing Ambivalence: A Multidimensional View of Attitudes Toward an Organizational Change.” Academy of Management Review, 25(4), 783–794. https://doi.org/10.5465/amr.2000.3707722

[24] Ford, J. D., Ford, L. W., and D’Amelio, A. (2008). “Resistance to Change: The Rest of the Story.” Academy of Management Review, 33(2), 362–377. https://doi.org/10.5465/amr.2008.31193235