I have watched a lot of organisations point to a human-in-the-loop control and call the risk handled.

The shape is always the same. A model or an agent produces something — a draft decision, a recommended action, a generated reply, a flagged transaction. A screen appears with two buttons. Someone clicks approve. Somewhere a log records that a human reviewed the output before it took effect. The control is now real on paper. The auditor can see it. The board paper can cite it. The deployment can proceed because a person was in the loop.

Then I ask the person who clicks the button three questions, and the control usually falls apart.

Could you have said no? Did you have time to look? Would you have known if it was wrong?

If the answer to any of those is uncertain, there was no control. There was a person absorbing accountability for a decision the machine had effectively already made. The loop was closed in the diagram and open in reality.

That is the work of D11, Workforce Literacy and Human-in-the-Loop Design, in the framework. It sits in the Enablement pillar, which is the pillar most likely to be read as soft — the training-and-change-management afterthought that gets a slide near the end. I want to argue the opposite. Literacy and HITL design are load-bearing controls. They are the controls that decide whether every gate the governance pillar designed actually holds when a real person is standing at it, tired, busy, and trusting the output more than they should.

The dimension is easy to flatten two ways: treat literacy as a completion metric — a percentage of staff who finished an awareness module — or treat human-in-the-loop as a universal safety blanket, a person in front of every model so the risk is contained. Both are comfortable and both are wrong. A literacy programme that nobody can connect to a better decision is a cost. A HITL point staffed by someone who cannot override it is a liability dressed as a safeguard.

The differentiator in D11 is D11.2, the HITL decision architecture: the deliberate design of where a human approves, where a human reviews after the fact, and where the system runs autonomously — and, underneath each of those, whether the human can actually exercise the authority the design assumes. This is not a training question; it is an operating-model question that connects directly to decisions made elsewhere. The autonomy levels and blast radius come from D5. The decision rights and accountability come from D6. The risk tiers and review paths come from D7. D11.2 is where those abstractions meet a named human in a named seat, and where the design has to answer whether that human is a control or a formality.

A genuine approval gate has three things behind it that a rubber stamp does not. The person has the authority to overrule the model without that override being treated as an error or a delay to be explained away. The person has the time the decision actually needs, rather than a throughput target that makes real scrutiny impossible. And the person has the competence to recognise when the output is wrong — the part that quietly fails most often, because the system tends to be confident and the human tends to be deferential. Strip any one away and the gate stops governing. Automation bias does the rest: people approve fluent, plausible output at a rate that rises as the system gets better, precisely when the residual errors get more subtle and more costly.

So the maturity question for D11.2 is not “is there a human in the loop”. It is “at which points did the organisation deliberately decide a human must be able to say no, and has it given that human the authority, the time, and the competence to mean it”. An automated credit-impacting decision, a clinical or safety-relevant recommendation, an employment-screening output, a customer-facing agent that can commit the business — these are not the places for a fast approve button next to a productivity target. A low-risk internal drafting assistant does not need the same gate at all. Designing the same human checkpoint everywhere is as much a failure as designing none: it trains people to click through the ones that matter because they have been clicking through the ones that don’t.

This is where the literacy in D11.1 stops being an awareness campaign and becomes the substrate the gates stand on. The framework’s literacy curriculum is tiered — executive, builder, and frontline — and the tiers exist because the decisions differ. The executive tier, which includes board AI literacy, is not about teaching directors to prompt; it is about whether the people accountable for the enterprise can read an AI risk posture, interrogate a value claim, and recognise when a control is theatre. The builder tier covers the people who design, configure, and evaluate AI systems, including the ones who design the HITL points themselves. The frontline tier covers the people who sit at the gates — the reviewers, approvers, and operators whose competence is the difference between a control and a checkbox.

I am deliberately not writing the curriculum. What such a programme should contain, and how adults actually acquire and retain operational competence rather than merely attend, is the discipline of instructional design and organisational change, with its own specialists. The operating-model question is not the syllabus. It is whether the literacy is tied to decision quality and to specific HITL roles, or whether it floats free as a generic completion number that proves nothing about whether the person at the gate can do the job the gate assumes. A literacy programme that cannot name which decisions it is meant to improve is the enablement-pillar version of D7’s well-organised hope.

D11.3 is the override and exception path, and it is where good intentions about human control most often die quietly. A reviewer who can approve but not effectively reject is not in control. If overriding the model is slow, career-risky, or treated as the reviewer second-guessing a validated system, people learn not to do it. The maturity signal is whether overrides happen, whether they are captured, and whether anyone looks at the pattern. A HITL design that never produces an override is not evidence that the model is always right. It is more often evidence that the gate is cosmetic. The exception path matters as much as the approval path: when the human escalates, does the escalation reach someone with the authority and competence to act, or does it dead-end? The override and exception flow connects back to the D6 decision rights and the D7 forum, because an override that cannot escalate to a real decision-maker is just friction.

D11.4 is the part the technology conversation tends to skip: the role redesign that AI forces and the organisation usually under-manages. When a model takes the first pass at work a person used to do end to end, the job changes shape. Some roles become supervision of machine output, which is a genuinely different skill — harder in some ways, because catching a confident machine’s subtle error demands more expertise than producing the first draft did, not less. Getting this wrong has two failure modes. One is hollowing out the very expertise the HITL gates depend on, so that in a few years there is nobody left who can tell when the machine is wrong, because the people who could have were the ones it replaced. The other is leaving people in redesigned roles without the recognition, the workload assumptions, or the support those roles now require. How that redesign is sequenced, how affected roles are consulted, and how any workforce-relations obligations are met is a specialist discipline — employment and industrial-relations counsel and organisational-change practitioners determine the requirements. The operating-model question is whether the role change is being designed deliberately, with the capability it depends on protected, or allowed to happen by accident as adoption spreads.

D11.5 is cultural readiness and psychological safety, which sounds the softest of all and is in fact a control. A reviewer who is afraid to flag a problem is a broken gate. If the organisation’s lived incentive is to ship fast and trust the AI, then the formal authority to override is fiction, because nobody will pay the social cost of using it. Psychological safety here is not a wellbeing nicety; it is the precondition for the override path in D11.3 to function. The maturity signal is whether challenge is observed and acted on — whether people actually raise concerns about AI output and whether raising them is rewarded or quietly punished. A culture that punishes the person who slowed the line to question a model has disabled its own most important control and will not find out until an incident reveals it.

There is a temptation, when a dimension touches workforce, to reach for hard regulatory citations to make it feel weighty. I want to resist that, because most of the AU regulatory force on D11 is indirect, and overstating it would do the dimension a disservice. There is no Australian statute that mandates an AI literacy curriculum or a particular HITL design for general enterprise. What is real is the accountability that flows from elsewhere and lands here. Where an entity is APRA-regulated, APRA CPS 230 makes operational-risk controls — including the controls a human-in-the-loop point is meant to be — something that has to actually work, not merely exist; specialist regulatory counsel determines the entity-specific obligations. Where a substantially automated decision affects an individual, the Privacy Act 1988 transparency provisions inserted by the Privacy and Other Legislation Amendment Act 2024, commencing 10 December 2026, shape what the organisation must be able to explain — which only sharpens whether the human in that loop understood the decision well enough for the explanation to be true; privacy counsel determines the specific application. For directors, the Corporations Act 2001 s180 care-and-diligence standard means an AI-related decision should rest on an informed basis, and board AI literacy under D11.1 is what makes “informed” more than a word. Beyond those, the Voluntary AI Safety Standard and the DISR AI Ethics Principles point at meaningful human oversight as direction of travel, voluntary rather than binding. The honest framing is that D11 is where other obligations become operable, not a dimension with a statute of its own.

This is also where “integrate enablement into existing committees” becomes too soft unless it is paired with hard accountability, the same failure I flagged in the governance dimensions. Literacy and HITL design need an owner with the mandate to fund capability uplift and the authority to make it a gate, not a recommendation. The natural owners sit across the CHRO and learning-and-development function, the COO who owns the operational gates, and the business-unit leaders who own the redesigned roles — but spread ownership without named accountability is the same as no ownership. Someone has to own whether the frontline tier of the literacy curriculum is actually current for the people sitting at material gates, and a deployment gate somewhere has to depend on it: a material AI use case should not pass its go-live review if the humans at its HITL points cannot demonstrate the competence the design assumes. The accountable path could run through a CAIO, an extended CIO/CDO/CTO mandate, a COO or CHRO-led enablement function, or CEO-direct sponsorship. The framework requires the accountable outcome and the gate, not a particular title — not capability uplift owned by everyone and therefore by no one, refreshed when there is time, and disconnected from any decision that depends on it.

The L4 buyer question for D11 is deliberately concrete: take one material AI use case with a human-in-the-loop control, and show that the human at that gate has the authority to override without penalty, the time the decision actually needs, and current role-specific literacy — and then show a real override from the last quarter that the system captured, escalated, and learned from. That question is hard to answer with a training-completion dashboard. It asks whether the loop is closed in reality and not just in the diagram. A figure would help here: a single HITL gate drawn with its three load-bearing supports — authority, time, competence — and the literacy tier feeding the seat, so the reader can see at a glance that removing any one support collapses the control.

So the test I would actually apply is small and quiet. Find the person who clicks approve on your most consequential AI-assisted decision. Ask them my three questions. Could you have said no. Did you have time to look. Would you have known if it was wrong. If the answers are yes, yes, and yes, then the enablement pillar is doing load-bearing work and the human in the loop is a control. If the answers wobble, then you do not have a human-in-the-loop safeguard. You have a person standing where a control was supposed to be, holding the accountability the machine handed them, hoping they are right.