Ask a Canadian firm how it decided what to automate and the answer has a predictable shape. Someone listed the tasks, assessed which ones a model handles well, and assigned those to the model. It is orderly, defensible, and rests on an assumption that the human factors discipline identified as false and named accordingly.

Key Takeaway

Function allocation by listing what humans and machines are respectively better at dates to Fitts in 1951 and carries an implicit assumption known as the substitution myth: that a machine can be inserted to do something a human is less good at without otherwise changing the nature of the work. Dekker and Woods established that inserting any technology fundamentally changes the functioning of the joint system, creating new tasks for the human such as entering inputs, engaging and disengaging the automation, and monitoring. The central irony is that these new tasks may require what the Fitts report itself said humans are bad at, namely tasks requiring vigilance and little activity. Researchers argue that quantitative "who does what" allocation does not work because the real effects of automation are qualitative: it transforms human practice and forces people to adapt their skills and routines. The constructive alternative is that automation is not binary. Sheridan and Verplank set out ten levels in 1978, and Parasuraman, Sheridan and Wickens associated those levels with distinct stages of information processing, so a level can be chosen independently for each stage of a task.

The Default Method

The method almost every organisation uses without knowing it has a name or a critique.

The procedure is to enumerate the work, judge each element against what the technology does well, and allocate accordingly. Data extraction goes to the machine. Judgment stays with the person. Classification goes to the machine. Client conversations stay with the person.

This is function allocation, which one source describes as a core activity of the human-machine systems discipline, and which Paul Fitts marked the outset of sixty years ago[1].

The reason it deserves examination is not that it produces obviously bad allocations. It usually produces sensible ones. The problem is that it answers a question that turns out not to be the operative one, and it does so in a way that systematically omits a category of cost. Both points are established in the literature and neither has reached the AI adoption conversation in Canadian professional services.

The Fitts List

The origin, and what it actually proposed.

Fitts published in 1951 on human engineering for an effective air navigation and traffic control system[2]. The approach became known as MABA-MABA, standing for Men-Are-Better-At and Machines-Are-Better-At lists[3].

Its logic is captured by one review: the assumptions are plausible because they capture the most important regularity of automation, namely that if the machine surpasses the human the function must be automated, and if not it does not make sense to automate. The list states that the primary, though not necessarily the only, driving force behind automation should be performance, meaning precision, power, speed and cost[1].

Sheridan called these the obvious advantages of automation, and Wickens explained the purpose of automation as performing functions the human operator cannot perform because of inherent limitations, performing functions the operator can do but performs poorly or at the cost of a high workload, and augmenting or assisting performance in areas where humans show limitations[1].

Two things worth noting before the critique. The framework is reasonable on its own terms and its persistence is not mere inertia. And Fitts himself was more cautious than the method's later users, predicting that for a good many years to come human beings would have intensive duties in relation to air navigation and traffic control[1], a prediction that has held for seventy-five years in one of the most automated fields there is.

Why It Persisted

A question worth asking, because the answer explains why finance functions reach for it instinctively.

The Fitts approach survives because it converts an intractable design question into a tractable list. It requires no theory of how humans and machines interact, only a comparison of capabilities on each element, and it produces an allocation that can be written down, approved and implemented.

It is also, in a narrow sense, correct. Where a machine genuinely surpasses a human at a well-defined function, and where nothing else changes, allocating it to the machine is the right call. The whole weight of the critique rests on the second clause.

The reason it appeals to a professional services firm specifically, and this is our own observation, is that professional work is already described in task terms. Engagements are decomposed into procedures, procedures into steps, and time is recorded against them. A firm therefore arrives at the automation question with a task inventory already built, and the Fitts method is the one that consumes that inventory directly.

That convenience is the trap. The inventory describes the work as it was performed manually, and the critique below is precisely that the work does not survive automation in that form.

The Substitution Myth

The assumption, named and stated.

One review sets it out directly: a more fundamental concern raised is the implicit assumption of MABA-MABA approaches that a machine can be inserted to do something that a human is less good at doing without otherwise changing the nature of the work, the so-called substitution myth[3].

The refutation is attributed to Dekker and Woods: inserting any technology into a system fundamentally changes the functioning of the joint system[3].

The word joint is carrying the argument. The unit of analysis is not the human and the machine considered separately but the system they form together, and a change to one component changes the system rather than swapping out a part.

The everyday version of the error is familiar once named. A firm calculates that a task took forty hours, that the system now performs it, and books forty hours saved. That calculation treats the forty hours as a removable block and the rest of the work as unaffected. The literature's claim is that neither is true: what remains has been altered, and new work has appeared that was not in the inventory.

The Tasks It Creates

The specific additions, which are enumerable and which almost no business case counts.

The same source states that new tasks are created for the human who now has to interact with the technology, giving examples of entering inputs, engaging and disengaging the automation, and monitoring[3].

Translated into a Canadian finance workflow, the created tasks include preparing and formatting inputs so the system can consume them; deciding when to invoke the system and when not to; monitoring outputs for correctness, which is the review control examined elsewhere in this series; handling the exceptions the system routes back; maintaining prompts, configurations and retrieval sources; and re-verifying after changes.

Several of these are substantial and recurring. Input preparation in particular has a way of migrating upstream: a system that requires clean, structured inputs imposes work on whoever produces them, which is frequently a different team whose costs do not appear in the automating function's business case.

The literature's list of consequences is broader still. One review notes that automation introduces problems including behavioural adaptation, mistrust and complacency, skill degradation, degraded situation awareness, problems when reclaiming control, and disruption to mental workload, and that automation changes the nature of human work often in ways unanticipated by designers[1].

That is six categories of consequence, none of which is a line in a productivity calculation, and several of which this publication has examined individually.

The Central Irony

The observation that turns the critique from a caution into a structural argument, and it is the single best sentence in this literature.

The review states that these new tasks, such as monitoring system states and functioning, may ironically require what the Fitts report originally stated humans are bad at doing, namely tasks requiring vigilance and little activity[3].

Follow the circle. The method allocates to machines the functions humans perform poorly. Doing so generates new human work consisting largely of monitoring. Monitoring is a vigilance task with little activity. And vigilance tasks with little activity are on the original list of things humans do badly.

So the procedure for removing work humans are bad at reliably produces work humans are bad at, and it does so as a structural consequence rather than as an implementation failure.

This is the same finding the deskilling article reached from Bainbridge and the automation bias article reached from Parasuraman, arriving here from function allocation theory. Three independent lines converge on the proposition that the monitoring role automation creates is a poor fit for human capability, and that is not a coincidence: they are describing one phenomenon from three angles.

For a Canadian finance function the design conclusion is that a workflow whose human component is pure monitoring has been allocated badly, regardless of how sensible the task-by-task reasoning was. The monitoring role should be a small part of the human's work rather than its substance.

Qualitative, Not Quantitative

The methodological conclusion the field drew, which reframes the whole question.

One paper argues that substitution-based function allocation methods such as MABA-MABA lists cannot provide progress on human-automation coordination, that quantitative "who does what" allocation does not work because the real effects of automation are qualitative, since it transforms human practice and forces people to adapt their skills and routines, and that rather than re-inventing or refining substitution-based methods, the more pressing question is how do we make them get along together[4].

Three claims sit in that passage and each matters.

That the effects are qualitative means they are changes in kind rather than in quantity. A firm measuring hours saved is measuring the quantitative dimension and missing the one the research says dominates.

That automation forces people to adapt their skills and routines means the adaptation is not optional and not a training matter. The work has changed, so the people change what they do, and they will do this whether or not anyone planned it.

And the reframing from who does what to how they get along relocates the design problem from allocation to interaction. The questions become how information passes between human and system, how the human establishes what the system did and why, how disagreement is handled, and how control is taken back.

Hollnagel's framing of moving from function allocation to function congruence[2] points the same way: fit between the parts rather than division between them.

The Uncounted Cost

The practical consequence for a business case, offered as our own analysis.

Assemble the argument. Automation removes some work and creates other work. The created work is largely monitoring, exception handling, input preparation and maintenance. The business case counts the removal and not the creation.

That is a sufficient explanation for a pattern many Canadian firms report: measured productivity gains from AI deployment that fall short of the projection, without anyone being able to identify where the difference went.

It went into work that had no line in the inventory because it did not exist when the inventory was built. Nobody was recording time against monitoring outputs, maintaining prompts or handling escalations, because none of those activities existed.

The remedy is procedural and cheap: before deployment, enumerate the tasks the deployment creates, estimate them, and include them in the case. The list from the literature is a workable starting point, being inputs, engaging and disengaging, and monitoring[3], extended with exception handling and maintenance.

A firm that does this will produce a smaller projected gain and a more accurate one, and will be in a position to notice when the created work grows.

Automation Is Not Binary

The constructive contribution, which is where this article turns from critique to method.

Sheridan and Verplank introduced in 1978 a scale of ten levels of automation, representing a continuum between low automation, in which the human performs the task manually, and full automation, in which the computer is fully automated[5].

The existence of a ten-point continuum matters because organisational decisions in this area are almost always taken as binary. A task is automated or it is not. A tool is adopted or it is not.

The intermediate levels describe genuinely different arrangements: the system suggests options while the human chooses; the system selects and the human approves; the system acts and informs the human; the system acts and informs only if asked. Each has a different risk profile, a different failure mode and a different effect on the human's engagement with the work.

Our observation is that most Canadian finance deployments cluster at two points on this scale, being a system that drafts for a human to accept, or a system that processes and reports. The intermediate arrangements, particularly those where the system surfaces options and reasoning while the human selects, are under-used, and they are the arrangements most consistent with the finding in the offloading literature that AI helps most when designed to question rather than to answer.

Levels Applied Per Stage

The refinement that makes the continuum usable, and the most practically important idea in this article.

Parasuraman took the Sheridan and Verplank scale and introduced the idea of associating levels of automation with functions, which are based on a four-stage model of human information processing and can be translated into equivalent system functions, so that the ten levels can be used to make distinct design choices for each of the four types of automation[5]. The model appears in Parasuraman, Sheridan and Wickens (2000)[2].

The unit of decision is therefore not the task. It is the stage within the task, and each stage can sit at a different level of automation.

That single move dissolves most of the false dilemma in AI adoption debates. The question stops being whether to automate a piece of work and becomes how much automation to apply to each stage of it, which permits a design in which the early stages are heavily automated and the later ones are not.

We should be precise about what we are reporting. Our source describes the model's structure, that it is based on a four-stage model of human information processing translated into system functions, without setting out the stage names, and we have not accessed the original paper. The mapping in the next section is therefore our own decomposition of a finance task, offered as an application of the principle rather than as a reproduction of the published taxonomy.

A Finance Decomposition

Applying the principle, with the stages framed for professional work. This section is our own analysis.

Consider a recurring finance task decomposed into four stages.

Gathering. Assembling the source material: documents, ledger extracts, prior files, correspondence. High automation is usually appropriate, since the failure mode is omission and is detectable by completeness checks.

Analysing. Extracting, classifying, computing, comparing. Moderate to high automation, with the important qualification that this is where faithfulness failures occur and where the outputs must remain traceable to sources.

Deciding. Selecting the treatment, forming the conclusion, choosing among defensible positions. This is where the level should generally be low, and where the intermediate arrangements matter most: a system that surfaces the options and the basis for each, with the human selecting, sits at a different level from one that presents a conclusion.

Executing. Posting, filing, issuing, communicating. The level here should reflect reversibility rather than difficulty. An irreversible or externally visible action warrants a lower level than an internal, correctable one, regardless of how capable the system is.

Two observations follow. A firm that decides "we automated the reconciliation" has made one decision where four were available, and has almost certainly applied a single level to stages with very different risk profiles.

And the reversibility criterion for the execution stage is not a capability judgment at all, which is precisely the point the Fitts method cannot express, since its only axis is who performs better.

The Static Problem

The limitation shared by every method described so far.

One source notes that methods based on the Fitts list, on levels of automation, or on cognitive task analysis share a common drawback, which is that the allocations derived from them are static, and observes that the relative strengths of humans and machines may not be static[6].

The point is straightforward and consequential. An allocation decided at design time encodes the relative capabilities as they stood at that moment, and both sides move: the system changes through provider updates and reconfiguration, and the human changes through the skill effects this series has documented.

An allocation that was correct at deployment can therefore become incorrect without anyone revisiting it, in either direction. The system may become capable of a stage that was reserved for the human, and the human may become less capable of a stage that was reserved for them.

For a Canadian firm the implication is that function allocation belongs on a review cycle rather than in a deployment document. The natural cadence is the one already argued for in the drift and incident articles, and the natural trigger is any material change to the system or the team.

Dynamic Allocation

The alternative the field developed, reported with its mechanisms.

One source records the suggestion that the issue and its attendant problems may be bypassed if function allocation is viewed as dynamic rather than static, so that allocation of function is not just a design activity but something that occurs during system operations, and notes interest in the effects of different forms of adaptive automation on human performance. It further reports that an early review identified five main categories of techniques for implementing adaptive automation: critical events, operator performance measurement, operator physiological assessment, modeling, and hybrid methods combining one or more of these[7].

Three of those five have plausible analogues in a finance setting, and this is our own mapping.

Critical events means the allocation shifts when a defined event occurs. In finance: material amounts, unusual counterparties, first-time transaction types, or period-end. The system's level drops and the human's role rises when the stakes rise.

Operator performance measurement means the allocation responds to how the human is performing. This is the same instrumentation the automation bias article recommended, repurposed from monitoring to control.

Modeling means an anticipatory model of workload or state drives the allocation.

Physiological assessment has no reasonable application in professional services and we mention it only for completeness.

The accessible version for a mid-sized Canadian firm is the first: a small set of rules that lower the automation level on defined triggers. That is materially better than a fixed allocation and requires no adaptive technology, only a decision written down in advance.

Which Function You Delegate Matters

An empirical result that adds a dimension the theory does not capture.

A study investigating whether classical function allocation holds for physical human-robot collaboration, using a within-subject design with 26 participants across four distinct allocations of position and force control in an abstract blending task, found that allocating position control to the human and force control to the robot, compared with the opposite, produced a significant improvement in preventing overblending, and was perceived better in terms of physical demand and overall system acceptance, with participants experiencing greater autonomy, more engagement and less frustration. The authors report a surprising insight that if position control was delegated to the robot, participants perceived much lower autonomy than when force control was delegated. They conclude that the findings empirically support applying Fitts' principles to static function allocation for physical collaboration while revealing nuanced user experience trade-offs, particularly regarding perceived autonomy[8].

We report this with its scope: a small sample, a physical manipulation task, and a domain remote from professional services.

The transferable finding is nonetheless notable. Two allocations delegating a comparable amount of the task produced markedly different experiences of autonomy, engagement and frustration. Which function is delegated matters independently of how much is delegated.

Our inference for finance work is that the stages differ in how much of a practitioner's sense of ownership they carry, and that the deciding stage is likely the one where delegation costs the most in perceived autonomy. If two designs offer similar performance, the one that retains the human in the stage they identify with is likely to be better received and better used, which connects to the resistance discussion in the deskilling article.

The Supervisory Surprise

A second finding from the same study that runs against this article's own argument, which is why it is worth reporting.

The authors record that the supervisory role, in which the robot controlled both functions, was rated second best in terms of subjective acceptance[8].

That is genuinely awkward for a straightforward reading of the monitoring critique. Everything above suggests the pure supervisory role should be least satisfying, and in this study it was rated second of four.

Two readings are available and we do not know which applies. Subjective acceptance may diverge from performance and from longer-run effects, so a role can be pleasant in a short study and corrosive over years, which is precisely the timescale on which deskilling operates. Or the monitoring critique may be overstated for tasks where supervision is genuinely light and the alternative is physically demanding.

We include it because reporting only the results that support the argument would be a form of the selection bias this publication has criticised elsewhere. The honest position is that the human factors literature's critique of the supervisory role is well established over long horizons, and that at least one recent empirical result found supervision more acceptable to participants than the framework would predict.

Ghost Work

A final category of created work, and one that operates at industry rather than firm level.

One review records that Gray and Suri debunked the myth of fully automated AI services, revealing the critical human role in tasks like image labeling and content moderation, describing this as ghost work, and noting that it underscores human creativity's importance and suggests automation creates new jobs while challenging traditional work cycles[9].

The relevance to the substitution myth is that it extends the argument beyond the deploying organisation. The human work created by automation is not only the monitoring and input preparation inside the firm; some of it sits in the supply chain that produced and maintains the system.

For a Canadian business the direct implication is modest, and we would not overstate it. But it does bear on a claim buyers hear frequently, which is that a service is fully automated. The research indicates such claims have generally been inaccurate about the systems they described, and a buyer is entitled to ask what human involvement exists in a service presented as autonomous, including where it sits and under what terms.

A Worked Case: The Same Amount, Two Designs

A Canadian firm automating a recurring analytical task. The reconstruction illustrates the design choices rather than reporting a specific engagement.

Design A follows the Fitts method. The firm lists the steps, judges that gathering, extraction, computation and drafting are all better performed by the system, and allocates them. The human reviews the finished output.

The result is a workflow whose entire human component is monitoring, which the literature identifies as a vigilance task with little activity and therefore as work humans are bad at[3]. The business case counted the four allocated steps as saved and did not count the review, the exception handling, the input preparation or the maintenance[3].

Design B applies levels per stage. Gathering is fully automated. Analysis is automated with outputs traceable to sources. The deciding stage is set low: the system presents the candidate treatments with the basis for each, and the practitioner selects. Execution is automated for reversible internal actions and requires approval for anything issued externally. A rule lowers the level further on material amounts and first-time transaction types.

Design B automates a similar proportion of the work. It differs in that the human's role includes a stage requiring active judgment rather than consisting solely of supervision, the deciding stage retains the ownership the autonomy finding suggests matters[8], and the created tasks were enumerated and costed.

The Fitts method could not have produced Design B, because its only axis is comparative capability, and three of Design B's choices, traceability, reversibility and the trigger rules, are not capability judgments at all.

What To Do

Stop deciding at the task level. Decompose into stages and choose an automation level for each. One decision where four were available is the most common design error.

Enumerate the created tasks before deployment. Inputs, engaging and disengaging, monitoring, exception handling, maintenance. Cost them into the case.

Refuse designs whose human component is only monitoring. That is a vigilance task with little activity, which is on the original list of things humans do badly.

Set the execution level by reversibility, not capability. Irreversible or externally visible actions warrant a lower level regardless of how well the system performs.

Use the intermediate levels. A system that surfaces options and reasoning for human selection is a different arrangement from one that presents a conclusion, and it is under-used.

Write trigger rules that lower the level. Material amounts, unusual counterparties, first-time transaction types, period-end. This is accessible dynamic allocation without adaptive technology.

Put allocation on a review cycle. Allocations are static and both sides move, so a correct allocation can silently become incorrect in either direction.

Consider which stage people identify with. Delegating comparable amounts produced markedly different experiences of autonomy, and the deciding stage is likely where delegation costs most.

Ask what human work sits inside a service described as fully automated. Such claims have historically been inaccurate about the systems they described.

The Limits Of This Analysis

Several caveats matter. The foundational works here, Fitts (1951), Sheridan and Verplank (1978), Parasuraman, Sheridan and Wickens (2000), and Dekker and Woods (2002), were all accessed through subsequent citation, review articles and repository records rather than in the original, and readers relying on any specific formulation should consult the primary sources. We describe the structure of the Parasuraman levels-by-stages model without reproducing its stage names, which our source did not set out, and the four-stage finance decomposition is our own construction offered as an application of the principle rather than as the published taxonomy. Two sources are personal or repository postings rather than peer-reviewed publications and are flagged. The empirical human-robot study involves 26 participants in a physical manipulation task, and its transfer to professional judgment work is our inference; we have reported one of its findings, on the acceptability of the supervisory role, that runs against this article's argument. The literature reviewed concerns aviation, process control and physical collaboration; the application to Canadian professional services throughout is our analysis rather than a demonstrated result. The uncounted cost argument, the finance stage decomposition, the reversibility criterion, the mapping of adaptive automation techniques to finance triggers and the worked case are our own. This article does not address workflow modelling methods, process mining, cost-benefit methodology, or organisational change management, which this publication addresses separately. Nothing here is a substitute for professional advice on workflow design in a specific environment.

Frequently Asked Questions

What is wrong with automating what AI is better at?
It carries the substitution myth: the assumption that a machine can be inserted to do something a human is less good at without otherwise changing the nature of the work. Dekker and Woods established that inserting any technology fundamentally changes the joint system, creating new tasks such as entering inputs, engaging and disengaging the automation, and monitoring.
What is the central irony?
That the new tasks created, particularly monitoring system states, may require what the Fitts report itself said humans are bad at, namely tasks requiring vigilance and little activity. The procedure for removing work humans perform poorly reliably produces work humans perform poorly, as a structural consequence rather than an implementation failure.
Why do measured gains fall short of projections?
A sufficient explanation is that the business case counts the work removed and not the work created. Monitoring, exception handling, input preparation and maintenance had no line in the task inventory because they did not exist when it was built. Enumerating and costing them before deployment produces a smaller projection and a more accurate one.
What should replace the Fitts method?
Levels applied per stage. Sheridan and Verplank set out ten levels between manual and fully automated, and Parasuraman associated those levels with distinct stages of information processing. So the unit of decision is the stage within a task, not the task, and different stages can sit at very different levels.
How should I set the level for the execution stage?
By reversibility rather than capability. An irreversible or externally visible action warrants a lower level than an internal correctable one, regardless of how well the system performs. That criterion is not a capability judgment at all, which is exactly what the Fitts method cannot express.
Does allocation need revisiting?
Yes. All these methods produce static allocations, and relative strengths are not static: the system changes through provider updates and reconfiguration, and the human changes through skill effects. An allocation correct at deployment can become incorrect in either direction without anyone revisiting it, so it belongs on a review cycle.
IB

About The Insight Bureau Research Desk

The Insight Bureau is GSH Financial's research publication, written for Canadian business owners and the students who will eventually advise them. This article reports an empirical finding that runs against its own argument, and marks where its practical framework is the authors' construction rather than a published taxonomy. See References below.

References

  1. Why the Fitts List Has Persisted Throughout the History of Function Allocation. Cognition, Technology & Work, on function allocation as a core discipline activity, the Fitts list's logic and its emphasis on performance meaning precision, power, speed and cost, Sheridan on the obvious advantages of automation, Wickens on the purpose of automation, the Fitts 1951 prediction about human duties, and the list of problems automation introduces. link.springer.com/article/10.1007/s10111-011-0188-1
  2. Designing Human-Automation Interaction: A New Level of Automation Taxonomy, repository record, for bibliographic references to Fitts (1951), Sheridan and Verplank (1978), Hollnagel (1999) on moving from function allocation to function congruence, and Parasuraman, Sheridan and Wickens (2000) in IEEE Transactions on Systems, Man, and Cybernetics. Note: accessed via a repository posting; the cited originals were not accessed directly. academia.edu/86677344
  3. Roth, E. M., Sushereba, C., Militello, L. G., Diiulio, J., & Ernst, K. (2019). Function Allocation Considerations in the Era of Human Autonomy Teaming. Journal of Cognitive Engineering and Decision Making, on the substitution myth as the implicit assumption of MABA-MABA approaches, Dekker and Woods (2002) on technology fundamentally changing the joint system, the new tasks created for the human, and the irony that these require what Fitts said humans are bad at. journals.sagepub.com/doi/full/10.1177/1555343419878038
  4. Human-Automation Interaction, repository record of work by Sheridan and Parasuraman, on the argument that substitution-based function allocation methods cannot provide progress, that quantitative allocation does not work because the real effects of automation are qualitative and transform human practice, and on reframing the question as how to make them get along together. Note: accessed via a repository posting. academia.edu/4177855
  5. Cerejo, J. (2021, February 1). Designing for Automation vs Augmentation, on Sheridan and Verplank's 1978 ten-level scale as a continuum from manual to fully automated, and on Parasuraman associating levels of automation with functions based on a four-stage model of human information processing, permitting distinct design choices per type. Note: a personal blog summarising the published models; the stage names are not set out in the source. medium.com/swlh/designing-for-automation-vs-augmentation-d2367a3ede42
  6. Function Allocation Considerations in the Era of Human Autonomy Teaming, repository record, on methods based on the Fitts list, levels of automation and cognitive task analysis sharing the drawback that their allocations are static, and on relative strengths of humans and machines not being static. researchgate.net/publication/337029165
  7. Human-Automation Interaction, repository record, on viewing function allocation as dynamic rather than static so that allocation occurs during system operations, and on five main categories of techniques for implementing adaptive automation: critical events, operator performance measurement, operator physiological assessment, modeling, and hybrid methods. researchgate.net/publication/240756917
  8. Mol, N., Prendergast, J. M., Abbink, D. A., & Peternel, L. Fitts' List Revisited: An Empirical Study on Function Allocation in a Two-Agent Physical Human-Robot Collaborative Position/Force Task. arXiv preprint 2505.04722, on the 26-participant within-subject design, the performance and acceptance advantage of allocating position control to the human, the perceived autonomy finding, and the supervisory role being rated second best in subjective acceptance. Note: preprint; a physical manipulation task with a small sample. arxiv.org/pdf/2505.04722
  9. Advancing Human-Machine Teaming: Concepts, Challenges, and Applications. arXiv preprint 2503.16518, on Roth and colleagues reviewing function allocation methods, and on Gray and Suri debunking the myth of fully automated AI services and revealing the critical human role described as ghost work. Note: preprint; the underlying work was not accessed directly. arxiv.org/pdf/2503.16518

This article discusses human factors and function allocation research and is provided for general informational purposes. The foundational works were accessed through subsequent citation and repository records rather than in the original, two sources are personal or repository postings rather than peer-reviewed publications, and the finance stage decomposition is the authors' construction rather than a published taxonomy. Nothing here is a substitute for professional advice on workflow design.