Skip to content

Version 1.0 · assessed 2026-08-10 · next review 2026-11-10

The AI Capability Ledger — what AI can and can’t do in the SME back office

This is the AI Capability Ledger: a dated, versioned record of what AI genuinely can and cannot do across an Australian SME's back office — bookkeeping, accounting, BAS, payroll, financial analysis, cash flow forecasting, IT support, AI agents, award interpretation and recruitment. Every entry carries a status, an honest limitation, and the basis it rests on. It is reviewed quarterly, and every change is logged.

It exists because the market has a truth problem. Vendors describe aspirations as capabilities, demos as products and pilots as proof, and the buyer is left to discover the gap after the contract is signed. An owner deciding whether to trust AI with their books, their pay runs or their hiring deserves a record that was written to be right, not to sell — including the entries that say 'not reliable' about things the industry is loudly claiming.

The statuses are deliberately blunt. 'Does well' means AI can reliably carry the bulk of that workload today, under routine human review. 'Assist only' means it is a genuine accelerant but its output is a first draft a person must judge before relying on it. 'Not reliable' means it produces confident errors often enough that it should not be trusted with the task. 'Human required by law' means an Australian legal reservation — not a technology gap — keeps the work with an accountable person, and the entry cites it.

Read it the way you would read a building inspection: the value is not in the reassuring lines, it is in knowing every line was assessed the same way. Where a capability is genuinely contested, we grade it down, and we say why. If you believe an entry is wrong in either direction, we want to hear it — corrections are how a ledger like this stays worth citing.

Does well× 45

AI can reliably carry the bulk of this workload today, under routine human review. Errors occur but are infrequent and catchable in a normal review layer; the economics of the task have genuinely changed.

Assist only× 14

AI is a genuine accelerant, but its output is a first draft. A person with the relevant competence must review and judge it before anything relies on it — either because errors are material, the capability is contested, or regulator guidance requires professional judgement to be applied.

Not reliable× 24

AI produces confident errors often enough, or lacks the context or judgement required, that it should not be trusted with this task. It may still assist around the edges, but the task itself belongs to a person.

Human required by law× 7

An Australian legal reservation — not a technology limitation — keeps this work with an accountable person: TPB registration for paid BAS and tax agent services, employer obligations under Fair Work and anti-discrimination law, or directors' duties under the Corporations Act. Each entry cites the specific reservation.

Finance

Bookkeeping

Full guide: Bookkeeping

Bank-feed transaction coding

Does well

AI learns how transactions have been coded before and suggests the account and GST code for each new bank line. For routine, repeating transactions — subscriptions, fuel, regular suppliers — suggestions are consistently usable and improve with more data.

Constraint: Reliable for the routine majority only; a person confirms, and unusual transactions must route to review.

Receipt and invoice data extraction

Does well

AI extracts supplier, date, amount and GST component from photographed receipts and emailed invoices into structured ledger data. One of the most mature applications — it removes most manual data entry from day-to-day bookkeeping.

Constraint: Extracted data still enters the review layer; the extraction is mature, the classification decisions downstream are not always.

Reconciliation matching

Does well

AI suggests matches between bank transactions and the invoices, bills and expense claims in the system, including harder cases like partial payments and combined deposits. A person confirms the match rather than hunting for it.

Constraint: Confirmation stays human; unmatched exceptions still need a person to resolve.

Anomaly and duplicate flagging

Does well

Having learnt the normal pattern of a ledger, AI flags what breaks it: duplicate invoices, unusually large payments, transactions coded inconsistently with their history. It surfaces questions a busy person would miss.

Constraint: A flag is a question, not a finding — a person decides what each flag means.

GST classification on routine purchases

Assist only

For common, well-understood purchases AI can suggest the GST treatment as it codes the transaction, which materially speeds up BAS preparation. It is a first draft for review, not a final answer.

Constraint: TPB(GS) 55/2026 requires AI output to be assessed and supplemented by professional judgement before being relied on.

Edge-case GST treatments

Not reliable

Mixed supplies, GST-free food line items, exports, insurance settlements, second-hand goods — these treatments depend on facts the software cannot see, and AI trained on routine transactions will confidently misclassify the unusual ones.

Constraint: Wrong classifications present with the same confidence as right ones; the TPB notes AI models may hallucinate or produce inaccurate information.

Owner and related-party transactions

Not reliable

Whether money drawn from a business is a wage, a loan, a repayment or a distribution is a judgement call with real tax consequences. AI sees a bank transfer; it cannot know the intent behind it or the structure it sits inside.

Constraint: These transactions need a person every time; the correct treatment depends on documents, intent and structure the ledger does not hold.

Capital versus expense classification

Not reliable

AI can guess from the description and amount whether a purchase is an operating cost or a capital asset, but the correct answer depends on what the item is for and how the business uses it — context that lives outside the ledger.

Constraint: Treatment changes tax outcomes; the deciding context is not in the transaction data.

Providing BAS services for a fee

Human required by law

Anyone who provides BAS services for a fee or other reward must be registered with the Tax Practitioners Board under the Tax Agent Services Act 2009. No AI tool holds a registration; a person or firm does, and answers to the ATO for the work.

Constraint: TPB registration is legally required for paid BAS services; TPB(GS) 55/2026 holds practitioners ultimately responsible for AI-assisted work.

Transaction processing at scale

Does well

Coding, matching and reconciling across high transaction volumes and multiple accounts is now largely automatable, with a consistency manual processing cannot match. This is the foundation layer under everything else in the function.

Constraint: Consistency at volume includes consistent errors — the review layer exists because the system cannot flag its own blind spots.

Accounts payable and receivable workflows

Does well

AI reads incoming supplier invoices, extracts the data, matches them against orders and payments, and flags mismatches and likely duplicates before money moves. On receivables it tracks who owes what and drafts follow-ups.

Constraint: People approve; AI prepares. Approval before money moves stays human.

Management-report assembly

Does well

Monthly profit and loss, cash summaries and management packs can be assembled by AI from a live ledger, so the compilation work stops consuming a qualified person's hours.

Constraint: Written commentary is a first draft only; a qualified person still decides what the numbers mean.

Variance and anomaly detection

Does well

AI compares the current period against the pattern of previous ones and flags what a human should look at: a cost line that jumped, revenue off trend, a margin quietly drifting. It turns 'review the accounts' into a short list.

Constraint: Judgement on whether a movement matters, and what to do about it, is human work.

Preparation for compliance work

Assist only

Assembling the figures behind a BAS, organising the records that support a tax return and reconciling clearing accounts before close — AI does the gathering and first-pass checking that used to be the slowest part of every compliance job.

Constraint: The output feeds positions a registered professional must review and stand behind; preparation is not the position.

Plain-language queries of your own ledger

Assist only

Modern tools answer questions like 'what did we spend with this supplier last quarter' directly from the ledger. Treated as a way to explore your own verified data, this is genuinely useful and low-risk.

Constraint: Answers should be sanity-checked against the source before anyone acts on them.

Related-party and unusual transactions

Not reliable

Transactions between connected entities, one-off events, disposals and settlements are reliably where automated treatment goes wrong: the ledger shows a movement, but the correct treatment depends on documents, intent and structure AI does not hold.

Constraint: AI does not reliably distinguish the routine from the exceptional — it processes both with equal confidence.

Tax positions, structuring and advice

Human required by law

Applying tax law to a business's specific circumstances, and advising on structure and timing, is professional judgement. Providing tax agent services for a fee or other reward requires registration with the Tax Practitioners Board.

Constraint: TPB(GS) 55/2026 is explicit that AI output is not a substitute for a practitioner's own analysis of a client's circumstances; TASA reserves paid tax agent services to registered practitioners.

Finance

BAS and GST

Full guide: BAS and GST

GST coding suggestions at volume

Assist only

For routine, well-understood transactions — the bulk of most businesses' activity — AI suggests the GST treatment as each transaction is coded, consistently and without fatigue, keeping treatments uniform across the period.

Constraint: Suggestions require review before reliance; TPB(GS) 55/2026 requires professional judgement to be applied to AI output.

Continuous reconciliation ahead of the BAS

Does well

AI keeps bank feeds matched to invoices and bills continuously, so pre-BAS reconciliation becomes confirmation of an already-current ledger rather than a quarterly archaeology exercise.

Constraint: Only as good as the bookkeeping discipline feeding it; exceptions still route to a person.

Pre-lodgement anomaly and duplicate detection

Does well

Before figures go near a statement, AI flags duplicated invoices, transactions coded inconsistently with their history, unusually large claims and GST codes that break the established pattern — a systematic version of the experienced pre-lodgement scan.

Constraint: Flags need a person to interpret; absence of flags is not evidence of correctness.

Draft BAS figure assembly

Does well

Pulling a period's coded, reconciled data into draft activity statement figures is deterministic, rule-following work AI and software automation do well. The draft arrives ready for review rather than being built by hand.

Constraint: The draft is an input to review, never the lodged statement.

Audit trail of classification decisions

Does well

Good AI-assisted workflows record what was classified, when, on what basis and who reviewed it — documentation that lines up with what the ATO expects of records and what the TPB expects of practitioners using AI.

Constraint: The ATO requires most business records to be kept for five years; records must be complete, retrievable and explainable regardless of the technology.

Edge-case GST treatments

Not reliable

Mixed supplies, GST-free food rules, exports, insurance recoveries, second-hand goods, property transactions — GST's hard cases turn on facts and definitions, not patterns, and AI processes the exceptional transaction with the same confidence as the routine one.

Constraint: A wrong GST code in an AI-assembled draft looks identical to a right one; only review separates them.

Adjustments versus corrections

Not reliable

Deciding whether a change belongs in this BAS as an adjustment or requires correcting an earlier statement is judgement about events, timing and thresholds — context that does not live in the transaction data.

Constraint: Where done for a fee this is registered-agent territory under TASA; either way the call is human.

Related-party transactions on the BAS

Not reliable

Movements between entities, or between an owner and their company, carry GST and tax consequences that depend on structure and intent. The ledger shows a transfer; the correct BAS treatment depends on what the transfer actually is.

Constraint: A person has to establish the character of the transaction before any treatment is defensible.

BAS review and lodgement for a fee

Human required by law

Providing BAS services for a fee or other reward requires registration with the Tax Practitioners Board. Whoever reviews and lodges the BAS stands behind it; AI never absorbs that role, it only changes how much preparation sits underneath it.

Constraint: TPB registration required under TASA 2009; TPB(GS) 55/2026 requires practitioners to assess AI output with professional judgement before relying on it.

Award-condition checking across every pay run

Assist only

AI can check rosters and timesheets against award conditions — penalty rates, overtime triggers, minimum engagement periods, allowances — for every employee in every pay run, rather than the sample a human reviewer has time to check.

Constraint: This is a checking layer under a compliance framework with human sign-off, not autonomous award interpretation; clause interaction and classification stay human calls.

Pay-run anomaly detection

Does well

By learning each employee's normal pay pattern, AI flags the pay that breaks it: a penalty loading that disappears, an allowance that stops, net pay that shifts without a roster change — exactly the errors that otherwise repeat for months.

Constraint: A person decides whether the anomaly is an error; the flag is the start of the check, not the end.

Leave-balance and entitlement tracking

Does well

AI tracks accruals continuously and flags balances that do not reconcile, negative balances building up, and long-service-leave thresholds approaching — the slow-moving entitlements manual processes notice late.

Constraint: Interpretation of entitlement rules on unusual employment histories still needs a person.

Classification-drift flagging

Does well

By comparing recorded duties, hours and responsibilities against classification indicators in the relevant award, AI can flag employees whose classification looks out of step — a common root of underpayment — and queue them for human review.

Constraint: Flagging only; resolving a classification is a judgement call with legal consequences and belongs to a person.

Super calculation and payday-super timing checks

Does well

With payday super live from 1 July 2026, AI can verify the 12% superannuation guarantee is calculated correctly on each pay event and that payments are on track to land within the seven-business-day window — every pay cycle, not once a quarter.

Constraint: AI tracks the clock; meeting the obligation, including awkward off-cycle and termination cases, remains the employer's responsibility.

Data hygiene and change monitoring

Does well

AI cross-checks new-starter details, tax file declarations and bank account changes, and flags unusual changes — such as a bank account edit just before pay day — that deserve a second human look.

Constraint: The second look must actually be human; account-change requests are a classic fraud vector.

Novel enterprise-agreement clause interpretation

Not reliable

Bespoke enterprise agreements contain wording no model was trained on. Interpreting an unusual clause — and deciding how it interacts with the award safety net — remains work for a person who can read intent, not just text.

Constraint: A plausible-sounding misreading propagates across every subsequent pay run unless a reviewer knows what wrong looks like.

Back-pay reconstruction

Assist only

When an underpayment is found, AI can assist the arithmetic of remediation, but reconstruction works from imperfect records, requires judgement on ambiguous periods and involves honest communication with affected staff.

Constraint: Humans must own the remediation process end to end; the calculation is the smallest part of it.

Resolving genuinely ambiguous classifications

Not reliable

Where an employee's duties straddle two classification levels, the right answer involves judgement, context and usually a conversation with the employee. AI can surface the ambiguity; it should not be left to resolve it.

Constraint: Misclassification is a common root of underpayment, and the employer wears the consequences.

Accountability for pay runs and records

Human required by law

'The software flagged it' is not a defence. Record-keeping, payslip and payment obligations sit with the employer, and a regulator's questions are answered by people, not tools.

Constraint: Fair Work requires time and wages records kept for 7 years and payslips within 1 working day of pay day; those obligations attach to the employer regardless of what technology runs payroll.

Finance

CFO and financial analysis

Full guide: CFO and financial analysis

Report and board-pack assembly

Does well

First drafts of monthly packs — statements, commentary skeleton, key movements — assembled from live data rather than copy-pasted from spreadsheets. The reviewing human edits and challenges instead of compiling.

Constraint: The pack cannot defend itself across a table; the presenting human's credibility is the asset.

Variance and trend detection

Does well

Machines are better than people at noticing that a number moved: margin creep, cost drift, a customer segment quietly shrinking — patterns spread across hundreds of transactions that no one reviewing totals would see.

Constraint: Detection is machine work; deciding whether the movement matters and what to do is not.

Drafted variance analysis and commentary

Assist only

AI compares actuals against budget and prior periods, isolates which lines drove the difference and explains it in plain English — the monthly 'why are we off?' conversation drafted before the meeting.

Constraint: AI can produce a fluent, well-formatted analysis containing a fabricated ratio or a misread account, and nothing in its tone will tell you; an experienced reviewer is the defence.

Scenario mechanics and stress-testing

Does well

Price rise, new hire, equipment purchase, losing a key customer — AI drafts multiple financial scenarios in minutes, each with its cash and margin implications laid out, so more scenarios actually get tested before a decision.

Constraint: The assumptions going in and the choice coming out are human; the model only reruns the arithmetic.

Continuous cash-runway monitoring

Does well

Instead of discovering the cash position at month end, AI watches it daily against commitments — payroll, super, tax, suppliers — and raises early flags when the runway shortens.

Constraint: Monitoring keeps the picture current; it does not discharge anyone's duty to act on it.

Ad-hoc financial question answering

Assist only

Questions that once meant an analyst and a day — software spend last year, which clients are less profitable than they look — now take a query against the ledger.

Constraint: Someone must sanity-check the answer before acting on it; fluency is not accuracy.

Judgement under ambiguity

Not reliable

Pricing decisions, when to raise capital, whether to take the contract with the difficult client — real CFO calls weigh incomplete information against competing goods. AI can lay out the options; it cannot weigh what matters most to this business and this owner.

Constraint: Strategic trade-offs are values questions wearing financial clothes; a model completes the pattern it was shown.

Banking and funding relationships

Not reliable

Lenders lend to people. Negotiating facilities, maintaining credibility through a rough quarter, making the case for headroom before you need it — relationship work no tool performs.

Constraint: Lenders and investors back people they can question across a table.

Keeping directors informed on solvency

Human required by law

Software can keep the financial picture current, but it cannot carry the duty. The human at the top of the finance function is how directors stay properly informed — and defensibly so.

Constraint: ASIC expects directors to be constantly aware of the company's financial position and to prevent an insolvent company trading (s588G Corporations Act); that duty lands on humans.

Finance

Cash flow forecasting

Full guide: Cash flow forecasting

Driver-based projections from live ledger data

Does well

Rather than extrapolating last month's bank balance, AI builds the forecast from underlying drivers — invoices raised, bills due, payroll cycles, recurring commitments — pulled straight from the accounting file.

Constraint: Bookkeeping quality is the ceiling on forecast quality: an unreconciled ledger produces a fluent, well-charted, wrong projection.

Receivables-timing patterns

Does well

AI learns each customer's real payment behaviour — who pays on time, who pays well past terms, who slips further in their own quiet season — and projects cash inflows from that evidence rather than from invoice terms.

Constraint: Learned behaviour lags reality; a customer wobbling towards insolvency is not in the pattern until it is too late.

Scenario stress-testing

Does well

What if the biggest customer pays a month late, or the hire starts in March instead of January? AI runs these branches in minutes, showing the cash consequence of each, so decisions get tested before they get made.

Constraint: The branches tested are the ones a human thinks to ask for.

Rolling refresh without rebuilds

Does well

The forecast updates as transactions land. There is no fortnightly rebuild to skip — which is what kills spreadsheet forecasts — and the model is only ever as stale as the bookkeeping behind it.

Constraint: Near-term numbers are working estimates, distant ones are direction; be wary of any tool that hides the widening range.

Seasonality detection

Does well

Given enough history, AI picks up the shape of the year — the December cliff, the quiet stretch after it, the end-of-financial-year invoice surge — and bakes it into the projection instead of leaving it to memory.

Constraint: Requires sufficient trading history; young or recently changed businesses do not have a pattern to learn.

Early-warning threshold alerts

Does well

The most useful output is not the chart but the alert: projected cash dipping below a threshold weeks out, while the options — chase debtors, defer spending, arrange facilities — are all still open.

Constraint: An alert nobody accountable is watching is not an early warning.

Forecasting through business regime change

Not reliable

When the business itself changes — new pricing, a new revenue line, a different customer mix — history stops being a guide, and a model trained on that history projects a business that no longer exists.

Constraint: Humans have to tell the forecast the world has changed; the model will not notice on its own.

Predicting one-off shocks

Not reliable

Losing your largest client, a flood, a supplier collapse — by definition these are not in the pattern. AI cannot predict them; what it can do is make their consequences fast to model once they happen.

Constraint: Fast consequence-modelling is valuable, but it is not foresight, and no honest tool claims otherwise.

Deciding the response to a projected gap

Not reliable

The forecast shows a gap; it does not decide what to do about it. Chasing debtors harder, delaying a hire, drawing a facility or having a frank supplier conversation are weighed against relationships and strategy.

Constraint: That weighing is owner-and-adviser work; material cash events also often live outside the ledger and must be fed in by a person.

Technology

IT support

Full guide: IT support

Around-the-clock first response

Does well

Every request gets an immediate, useful reply at any hour, including a clear answer on what happens next. Nothing sits unread over a weekend, and urgent items are flagged to a person straight away.

Constraint: First response is not resolution; escalation paths to humans must actually exist and be staffed.

Triage, classification and routing

Does well

AI reads the request, classifies it, attaches the relevant history and sends it to the right queue. Escalations arrive with context instead of a one-line subject, which shortens the human part of the fix.

Constraint: Misrouted edge cases still occur; a person owns the queue.

Password resets and account unlocks

Does well

One of the most common ticket categories, handled through verified self-service flows with identity checks built in — the AI guides the process; the security controls decide who gets back in.

Constraint: The identity verification is done by the security controls, not by the AI's judgement.

How-do-I answers from your own documentation

Does well

Questions about software you already run — sharing a calendar, setting an out-of-office, connecting a printer — answered from your environment's documentation rather than generic web advice.

Constraint: Quality is bounded by the documentation; undocumented environments get generic answers.

Known-issue runbook fixes

Does well

Where an issue has a documented, low-risk resolution — clearing a stuck update, re-mapping a drive, reinstalling an application — AI can execute the runbook and record exactly what it did.

Constraint: Documented and low-risk only; anything novel or destructive is out of scope by design.

Pattern spotting across tickets

Does well

AI notices when the same fault keeps recurring across different people — a failing update, a flaky switch, a licence about to lapse — and turns repeat symptoms into one root-cause job for an engineer.

Constraint: The root-cause work itself is engineering, not pattern matching.

Novel incident resolution

Not reliable

AI resolves what has been seen and documented before. The outage nobody has seen — the strange interaction between an update and your line-of-business software — needs a person who can reason from first principles.

Constraint: Pattern matching fails precisely when the incident has no pattern.

Security judgement calls

Not reliable

Deciding whether an alert is noise or the start of an incident, whether to isolate a machine, whether anyone needs to be notified — judgement calls with real consequences. AI can surface the signals; a human must make the call.

Constraint: ACSC guidance for small business assumes human oversight of AI use; incident decisions carry consequences no tool answers for.

Requests touching money or identity

Not reliable

Changing bank details, granting new access, elevating permissions, processing an urgent payment request from the director — these are precisely the requests attackers forge.

Constraint: Always require human verification, no matter how confident the AI sounds.

Physical and network work

Not reliable

Cabling, hardware failures, Wi-Fi dead spots, a new office fit-out — AI cannot hold a screwdriver. A service that leans entirely on AI has quietly excluded a large part of what IT support actually is.

Constraint: Physical work is simply outside the medium; this is a hard boundary, not a maturity gap.

Technology

AI agents and automation

Full guide: AI agents and automation

Drafting and summarising

Assist only

First drafts of emails, quotes, meeting notes and reports, assembled from your own documents and data. The economics are simple: the agent does the assembling; a person does the deciding.

Constraint: Drafts leave the business only after a human decision; generated text can be fluently wrong.

Inbox and request triage

Does well

Reading incoming email or form submissions, classifying them, extracting the details that matter and routing them to the right person with a summary attached — reliably, at any hour.

Constraint: Routing errors happen at the margins; a person owns the queue and its exceptions.

Moving data between systems

Does well

Invoices into accounting software, applicants into a spreadsheet, orders into a job list — bridging work that is repetitive but slightly variable each time: too messy for rigid automation, too dull for a person.

Constraint: Chains break at the joins — an expired login, a changed invoice layout — and can fail silently for weeks without monitoring; someone has to own the checking.

Monitoring and exception flagging

Does well

Watching for the thing that should not happen: an invoice past terms, stock below reorder, a review awaiting a reply, a certificate about to expire. Agents are tireless at surveillance work people do inconsistently.

Constraint: The agent raises the flag; acting on it stays with a person.

Research and comparison legwork

Assist only

Gathering supplier options, summarising documents, lining up quotes in one table. The agent compiles; you judge.

Constraint: Treat the output as a well-organised starting point, not a verdict.

Customer FAQs within documented limits

Does well

Answering the questions that have documented answers — opening hours, pricing basics, order status — and handing everything else to a person. The honest version knows what it does not know.

Constraint: Scope discipline is the safety mechanism; an agent answering beyond its documentation is guessing in your brand's voice.

Unsupervised consequential actions

Not reliable

An agent that can send, pay, post or delete without review will eventually do one of those things wrongly — and confidently. Over-broad permissions turn a small error, or a manipulated instruction, into a large one.

Constraint: Keep a human approval step on any action that leaves the business or changes a record you would have to explain later; least privilege applies to software colleagues too.

Handling personal information in agent workflows

Assist only

Whatever an agent can read, it can mishandle: client records pasted into consumer AI tools, or an over-broad connection to file storage, can put sensitive data somewhere you cannot retrieve it. Decide what the agent may see before deciding what it may do.

Constraint: The OAIC is clear the Privacy Act applies to all uses of AI involving personal information, and recommends against entering personal — especially sensitive — information into publicly available generative AI tools.

People

Award interpretation

Full guide: Award interpretation

Cross-checking pay runs against published rates

Does well

After a pay run is calculated, AI can compare outcomes against currently published minimums and flag anything sitting below or oddly against them. Checking a specific number against a specific source is the shape of task AI does reliably.

Constraint: It is generating rates from memory that fails, not checking against a supplied source; the Fair Work Ombudsman's Pay and Conditions Tool remains the authoritative reference.

Catching annual wage review changes

Does well

Award minimums move after every Annual Wage Review — from 1 July 2026, award minimum rates rose 4.75%. AI monitoring is well suited to noticing that a rate in payroll settings no longer matches a published one, before an underpayment compounds.

Constraint: Monitoring detects the mismatch; applying the correct new rate and back-checking affected periods is human work.

Plain-English clause summaries with the source supplied

Assist only

AI is genuinely good at turning a clause into a readable explanation — provided the actual clause is supplied and cited, so a person can verify against the source rather than trusting a paraphrase drawn from training data.

Constraint: Without the source clause attached, a summary is an unverifiable paraphrase and should not be relied on.

Anomaly flagging across the roster

Does well

Patterns a person checks occasionally, AI checks every pay cycle: unusual overtime, allowances that quietly stopped appearing, penalty-rate hours that do not line up with rostered times.

Constraint: It raises the flag; a person decides what the flag means.

Preparing the question for a human adviser

Does well

When something genuinely needs an adviser — or the Fair Work Ombudsman — AI can assemble the picture first: the classification in use, the hours, the clauses in play, the discrepancy that triggered the review. Better questions in, faster and cheaper answers back.

Constraint: Assembly of the facts, not resolution of the question.

Answering award questions from memory

Not reliable

Ask a general chatbot a specific award question and it will answer confidently and often wrongly. Superseded rates, invented clause references and mixed-up award versions all present identically to the right answer.

Constraint: Models are trained on the past and award rates change every July; fluency is not accuracy.

Multi-clause interaction calculations

Not reliable

Award questions are rarely one clause. Casual loading, penalty rates, overtime, allowances and enterprise-agreement overlays interact, and the order of operations changes the answer. General AI does not reliably execute that interaction; it approximates it, plausibly.

Constraint: A plausible approximation of a pay calculation is an underpayment risk, not an answer.

Classification judgement

Not reliable

Whether someone is a Level 2 or a Level 3 depends on the duties actually performed, not the title in the contract. That is a judgement call with legal consequences, made by a person who understands the work.

Constraint: An AI has only the words it was given; the duties actually performed live outside the text.

Carrying liability for pay outcomes

Human required by law

If AI gets payroll wrong, the employer wears the underpayment, the back-pay and the compliance consequences. No AI vendor stands behind an award interpretation the way an accountable payroll provider must.

Constraint: Fair Work record-keeping (7 years) and payslip (1 working day) obligations sit with the employer, not the tool; the Fair Work Commission has stated it will not use generative AI to make decisions under the Fair Work Act.

People

Recruitment and screening

Full guide: Recruitment and screening

Job ad drafting and placement

Does well

A solid ad tailored to role and channel now takes minutes, not a morning, and distribution across boards and networks is push-button.

Constraint: The craft left in ads is knowing what will actually attract the person you want — that judgement stays human.

Sourcing and longlisting

Does well

Searching databases and networks for plausible candidates was the recruiter's protected asset. AI matching does the first pass wider and faster than any human researcher.

Constraint: A longlist is a starting universe, not a shortlist; matching finds plausibility, not suitability.

Screening at volume

Assist only

Parsing hundreds of applications against criteria and suggesting a shortlist is machine work now — the reading-every-CV grind is over.

Constraint: Screening criteria can encode bias, and AHRC guidance puts the onus on employers to ensure fair, lawful hiring practices; a human must check what the criteria actually select for.

Scheduling and coordination

Does well

The email tennis of getting three calendars to agree, rescheduling and reminders — fully automatable, and one of the largest hidden time costs in any hiring round.

Constraint: Logistics only — the interview the schedule serves, and the judgement inside it, stay human.

Candidate communication at scale

Does well

Acknowledgements, status updates and polite declines can now go to every applicant, promptly. Automation fixes recruitment's worst reputation problem: silence.

Constraint: The sensitive conversations still belong to people.

Judging candidate fit

Not reliable

Interviews, reference conversations, the read on whether someone will thrive in this team under this manager — hiring is a judgement about a specific human in a specific context, made on incomplete information.

Constraint: This is the least automatable task in the entire back office.

Handling candidate personal information

Assist only

CVs and interview notes are personal information, and handling candidate data properly is now part of hiring competence. AI can process it inside properly configured systems.

Constraint: The OAIC is clear the Privacy Act applies to all uses of AI involving personal information, and recommends against putting personal — especially sensitive — information into publicly available generative AI tools.

Selling the role and closing candidates

Not reliable

The best candidate usually has options. Understanding what they actually want, positioning the role honestly, and steering an offer through negotiation and cold feet is persuasion built on trust.

Constraint: A human trade, and the one that decides whether the whole funnel produced anything.

Accountability for a fair hiring process

Human required by law

Anti-discrimination law applies to hiring however the decision is produced, and AI screening needs a human checking what its criteria select for. The tool never carries the liability; the employer does.

Constraint: AHRC guidance puts the onus on employers to ensure whoever recruits for them — staff or agent — knows their legal obligations under Australian anti-discrimination law.

Methodology

How this ledger is assessed

Version 1.0 is a consolidation, and we want to be precise about what that means. Every entry is drawn from the verified research behind the 16 AI pages Valont published in August 2026: Australian regulator guidance checked against primary sources, vendor and platform documentation reviewed during that research, and Valont's own production use of AI across client back-office work. It is an assessment, not a laboratory benchmark — we have not run controlled tests on named products, and no entry claims we have.

Each claim traces to a documented research trail. Regulator positions — the TPB's guidance on AI use TPB(GS) 55/2026, ATO record-keeping rules, Fair Work record-keeping and payslip obligations, ASIC's directors' duties guidance, OAIC guidance on commercially available AI products, ACSC small-business AI guidance and AHRC recruitment guidance — were fetch-verified against the primary source on 10 August 2026, with the verified wording recorded in the source trail behind each page. Capability observations come from that same reviewed corpus; where a claim lacked a traceable basis, it was left out of the ledger rather than smoothed over.

Two disciplines govern the grading. First, no invented figures: the ledger quotes no accuracy percentages or error rates, because none were verified — a number would look more rigorous and be less honest. Second, contested capabilities are graded down: where reasonable practitioners disagree about whether AI 'does well', the entry says assist-only and the detail explains the doubt. The 'human required by law' status is used only where a specific Australian legal reservation exists, and the entry cites it.

The ledger is reviewed quarterly. At each review, every status is re-assessed against what has actually changed — regulator guidance, platform capability, and what we see in production use — and every movement is recorded in the change log with a date and a reason. Between reviews, corrections are invited: if a vendor, practitioner or regulator can show an entry is wrong, we will change it and log the change. The record of revisions is part of the asset, not an embarrassment to it.

Change log

  • v1.0 · 2026-08-10Initial ledger: consolidation of the verified research behind the 16 AI pages published August 2026.

FAQ

Frequently asked questions

Can't find the answer you're looking for? Get in touch

Want this working in your business, not just on a page?

The ledger shows what AI genuinely does well under human review. A connected back office is where that actually happens — one team, one shared picture of the business, judgement where it matters.