The First Hour Costs the Most: Per-Visit IT Billing Against a Retainer
A method for the finance or procurement lead holding two IT quotes shaped differently — one per hour, one per user per month. Why they cannot be compared as written, how each model actually bills, a five-step conversion onto one axis, the three risks that appear in neither number, when per-visit is genuinely right and when it quietly stops being, and the hybrid almost every organisation lands on.
Published
The short answer: a per-visit quote and a retainer quote cannot be compared as written, because they are not priced on the same axis. Per-visit prices an *event*. A retainer prices *availability*. Convert both into cost per month at your real ticket volume, then look separately at the three risks that never appear in either number.
Two quotes, two units, no way to compare them
A finance or procurement lead ends up holding two pieces of paper. One says something like "US$98 for the first hour on site, US$78 for each additional hour". The other says something like "HK$1,247.40 per user per month". Both providers are credible. Both quotes are honest. And there is no arithmetic on either page that turns one into the other.
This is not a trick, and neither provider is hiding anything. The two documents are selling different units. The first sells an event: something breaks, somebody comes, you pay for the time they were there. The second sells availability: a service exists, it is staffed, it is watching, and it costs the same this month whether you used it heavily or not at all.
Comparing them as written produces a predictable mistake. The hourly quote always looks cheaper on the day you read it, because on that day the number is zero. The retainer always looks expensive, because the number is non-zero before anything has happened. Finance teams that buy on that impression tend to discover the real cost eighteen months later, in a year-on-year variance they cannot explain.
The fix is not complicated. Put both on the same axis — cost per month at your actual volume — and then, separately, look at the three things that neither number contains.
How per-visit billing actually works
Per-visit billing has a shape, and the shape matters more than the headline rate.
The first hour costs more than the hours after it. Read from Brocent's dispatch and on-site rates page on 22 September 2026, Hong Kong is listed at US$98 for the first hour and US$78 for each additional hour; Singapore at US$85 and US$78; mainland China at US$65 and US$59. The basis line is part of the price: Level 1 end-user-computing support, next-business-day, standard 9×5, indicative 2026 H2 figures, US dollars, tax exclusive, city centre, falling at higher volumes, and fully loaded — which here means the rate absorbs cost of living, bilingual engineers, training, FX and tax handling, credit-term financing and 24×7 coordination rather than adding them later as line items.
The reason the first hour is priced higher is not margin. It is travel, scheduling, the engineer's lost slot either side of your job, and the fixed cost of turning up at all. Which leads to the single most important sentence in this article: a twenty-minute job is never a twenty-minute charge. It is a first hour. Under prepaid tokens, where the token service bills a two-token minimum for a business-hours on-site visit, it is two hours. Under any per-event model anywhere, it is a minimum.
That fact quietly destroys most naive comparisons. A site with lots of small, physical, frequent problems is the worst possible fit for per-visit billing, because almost every charge is a minimum charge, and the average cost per useful minute is enormous. A site with occasional substantial jobs — a half-day of cabling, a morning of new-starter setups — is the best possible fit, because the expensive first hour is amortised across a long visit.
So: work out what your visits look like, not just how many there are.
- A four-hour visit in Hong Kong bills US$98 plus three hours at US$78 — US$332, an average of US$83 per hour.
- Four separate one-hour visits in the same month bill US$392, an average of US$98 per hour.
Same four hours of engineering, eighteen percent apart, decided entirely by whether the work arrived batched or scattered.
Two further mechanics are worth reading in any per-visit contract, because they are where unmodelled cost hides. Travel is often loaded separately for sites outside the metro area — under the published token billing rules, a remote site adds one token under 30 km and two tokens between 30 and 50 km, capped at two per visit. And out-of-hours work carries a multiplier: the same published token rules set a two-token minimum in business hours, three in the evening, four late at night or at weekends, and six on public holidays. If a meaningful share of your physical work happens outside 9×5 — and in retail, manufacturing and trading operations it usually does — a daytime headline rate is not the rate you will pay.
How a retainer actually works
A retainer prices availability, and it is usually denominated per user or per device per month.
On Brocent's published Hong Kong per-user plans, the tiers run at HK$855.14 per user per month for the smallest band (1–5 employees), HK$1,247.40 for the 5–300 band and HK$1,561.21 for the 10–500 band, with an enterprise tier priced on application. A thirty-person Hong Kong office on the middle band is therefore looking at roughly HK$37,400 a month — a number that does not move when a quiet month follows a busy one.
What that buys is a standing function rather than a set of visits: the service desk and its coverage window, patching, endpoint security, identity administration, backup and its verification, monitoring, documentation, and a defined escalation path. It is the difference between having somebody to call and having somebody already watching.
Two things determine whether a retainer is good value, and both are contractual rather than numerical.
What is in and what is out. Retainers are not unlimited. Read the scope for the boundary between "covered" and "project work": migrations, office moves, new-site builds, hardware procurement and after-hours work are commonly outside the per-user fee and quoted separately. That is reasonable, but it needs to be visible before signing, not discovered in month four.
Where the overage sits. The honest question to ask a retainer provider is what happens when consumption runs far above what the pricing assumed. Good answers describe a mechanism — a fair-use boundary, a review trigger, a rate for excess on-site hours. Bad answers say "we don't really track that", which means either the price already assumes low usage or someone will be having an uncomfortable conversation later.
Putting them on one axis
Here is the method. It takes an afternoon and it is worth more than any vendor comparison table.
Step one: count the events, not the tickets. Pull twelve months of history and isolate the items that actually required a person on site. Most organisations find this is a small fraction of total tickets, and the number surprises them.
Step two: work out the shape. For each on-site event, record roughly how long it took and whether it could have waited a week. You now know your monthly visit count, your average visit length, and — critically — what share of your work is batchable.
Step three: price the per-visit column. Multiply out at the published rates for your market, using first-hour-plus-additional-hours arithmetic rather than a flat average. Add travel loading for any site outside the metro. Add the out-of-hours multiplier for the share that falls outside 9×5. The result is a monthly figure.
Step four: price the retainer column. Per-user rate times headcount, plus anything your scope reading identified as excluded that you would nonetheless need.
Step five — and this is the step everyone skips: notice what is missing from the per-visit column. The retainer figure includes patching, monitoring, backup verification, identity administration and a service desk. The per-visit figure includes none of these. If you are comparing them directly you are comparing a whole service against a subset of one, and the per-visit column will win a race it was never actually running in.
Done properly, step five usually reframes the question entirely. The real comparison is almost never "per-visit or retainer". It is "retainer alone" against "retainer plus a per-visit or prepaid arrangement for the physical layer" — which is a question about how to cover hands, not about whether to buy a service.
Per-visit dispatch vs a block of prepaid hours vs a monthly retainer
- Per-visit dispatch — unit of purchase: one visit. What it guarantees: a response governed by the SLA tier you bought — published tiers run from a two-hour P1 first response with next-business-day arrival, through one hour and same-day, to fifteen minutes and four-hour arrival on the 24×7 emergency tier. How it fails: frequency. Every small job pays a minimum, and no one accumulates knowledge of your site. Invoiced after the job, itemised by ticket, on standard 30-day terms.
- Block of prepaid hours — unit of purchase: a pack of tokens, one token equalling one hour of on-site or remote work, in packs of 20, 50, 100 or 250. What it guarantees: a rate locked in advance and a balance that does not expire for twelve months, with a one-time goodwill extension available each contract year. Effective hourly rate is simply pack price divided by tokens, so larger packs are cheaper per hour — in Hong Kong the published packs work out at roughly HK$391 per hour at 20 tokens, HK$348 at 50 and HK$304 at 100. How it fails: the minimums still apply — two tokens for a business-hours on-site visit, more in the evening, at weekends and on public holidays — and a balance bought for a busy year becomes a sunk cost in a quiet one.
- Monthly retainer — unit of purchase: a user (or a device), per month. What it guarantees: a standing service — desk, patching, security, identity, backup, monitoring, documentation, escalation — that is staffed whether or not you call. How it fails: scope drift and low utilisation. A retainer bought by an organisation that genuinely has almost no IT need is an expensive insurance policy; one bought by an organisation that has a lot of need is the cheapest thing on this list per unit of work.
The three risks that appear in neither number
Price is the easy part of this decision. Three risks sit outside both quotes, and in practice they decide more outcomes than the arithmetic does.
Response-time risk. Per-visit models respond according to the tier you bought; if you bought the cheapest tier, "next business day" is the commitment, and on a Friday afternoon that means Monday. A retainer usually carries a defined priority framework with response and resolution targets per severity. If a four-hour outage genuinely costs your business more than a month of retainer, the cheaper model is the more expensive one. Our guide to SLA priority levels sets out how P1 to P4 targets are normally defined and measured.
Availability risk. Per-visit providers allocate engineers to whoever is in the queue. You are not reserved. In a genuinely busy week — the week of a regional outage, or the week before Chinese New Year — an unreserved buyer waits behind reserved ones. A retainer is, among other things, a claim on capacity. That is a real part of what the monthly fee buys and it is invisible in the rate card.
Knowledge risk. This one compounds silently and is the most expensive of the three over a three-year horizon. Under per-visit billing, whoever turns up is whoever was free. They will spend the first part of every visit rediscovering your environment — where the comms cabinet is, what the Wi-Fi is called, which switch feeds the meeting rooms, why that one server is not on the domain. You pay for that rediscovery every time, at first-hour rates, and it never accumulates into anything. A retained provider builds a configuration record and runbooks, so visit number twelve is not a repeat of visit number one. If your environment is at all non-standard, this is usually the deciding argument.
When per-visit is genuinely right — and when it quietly stops being
Per-visit is the correct buy more often than retainer-selling providers admit. It is right for a satellite site with a handful of users and no server room. It is right for an office whose real support load is overwhelmingly remote, with a physical trickle. It is right when demand is genuinely unpredictable and low, because you pay for spikes only when they occur. And it is right as a *measurement instrument*: six months of itemised per-visit invoices is the best sizing data anyone can hand you, because it records what actually happened rather than what someone estimated.
It stops being right when three things happen together, and they tend to arrive at once. Volume becomes steady rather than spiky. The jobs get small, so almost every charge is a minimum. And the environment grows enough that rediscovery time per visit starts to dominate the actual work. At that point the per-visit invoice is still arriving in small pieces, which is why nobody notices, but the annual total has quietly passed what a retainer would have cost — and the organisation has nothing to show for it: no documentation, no monitoring, no patch baseline, no one who knows the place.
The practical test is to total twelve months of per-visit invoices and divide by twelve. Most organisations have never done this, because the invoices arrive separately and are approved separately. Doing it once is usually decisive.
The hybrid almost everyone lands on
The two models are not rivals in practice, because they cover different layers.
The retainer covers the remote layer: the desk, the patching, the security posture, the identity administration, the backup and its verification, the monitoring, and the documentation that makes every subsequent visit shorter. This layer is also what shrinks the physical pile in the first place — a well-patched, well-monitored estate simply generates fewer incidents that need hands.
The per-visit or prepaid arrangement covers the hands layer: the things that have to happen in the room. Sized off the ticket history rather than a guess, this layer is usually far smaller than the organisation expected before it counted.
Buy them in that order. The remote layer first, because it determines how big the physical layer needs to be; the physical layer second, sized to what is left. An organisation that buys the hands layer alone ends up reconstructing the remote layer by accident — a backup tool here, an antivirus subscription there, a spreadsheet of passwords somewhere — and pays more for a worse version of it.
Brocent publishes the inputs for both sides of this arithmetic: the per-visit and dedicated-engineer rates on the dispatch and on-site rates page, prepaid packs and their billing rules on the token service page, and per-user plans on the pricing pages. One note on reading them together: on-site dispatch is quoted in US dollars, while token packs and per-user plans are quoted in local currency. They are not directly convertible into a like-for-like hourly comparison without fixing the visit shape, the coverage window and the SLA first — so compare within a model, and let the method above rather than a currency conversion decide between models.
Frequently asked questions
Why does the first hour cost more?
Because turning up has a fixed cost that has nothing to do with the length of the job: travel, scheduling, and the slot lost either side of your visit. The first hour absorbs it; subsequent hours do not, which is why the additional-hour rate is lower. The practical consequence is that batching work into fewer, longer visits is materially cheaper than scattering it across many short ones.
Does travel get billed separately?
It depends on where the site is. Within the metro area, travel is normally inside the first-hour rate. Outside it, a loading usually applies — under the published prepaid token rules, one extra token for a site under 30 km away and two for 30 to 50 km, capped at two per visit. Ask specifically where the boundary sits for your address rather than assuming.
What is a fair minimum purchase?
For prepaid packs, the published minimum first purchase is twenty tokens per service zone, which is also the smallest pack. For per-visit dispatch there is no monthly minimum at all — you can buy a single visit. What to watch instead is the *per-visit* minimum: a business-hours on-site visit bills at least two tokens, and more in the evening, at weekends or on public holidays.
Can unused retainer hours roll over?
A per-user retainer does not usually contain "hours" to roll over — it buys a standing service, not a quantity. Prepaid blocks are the model that carries a balance, and there the relevant term is validity: published tokens are valid twelve months from purchase, with a one-time goodwill extension available each contract year through your account manager. If a provider offers rolling hours inside a retainer, read carefully what happens to them at renewal.
What happens when we exceed the block?
You top up, and the rate you pay for the top-up is the question to settle in advance. Larger packs carry lower effective hourly rates — pack price divided by tokens — so a mid-year top-up bought as a small pack is more expensive per hour than the original purchase. If your consumption is trending above plan, it is usually cheaper to move up a pack size than to keep buying small ones.
Which model gives us a real SLA?
Both can, but they express it differently. Per-visit SLAs are bought as a tier and govern response and arrival — published tiers run from a two-hour P1 first response with next-business-day arrival up to fifteen minutes and a four-hour arrival on the 24×7 emergency tier. Retainer SLAs are usually expressed as a priority framework with response and resolution targets by severity. The question to ask either provider is not "do you have an SLA" but "what happens contractually when you miss it".
Which is cheaper for two sites?
Usually the retainer, and the gap widens with distance. Two sites double the travel exposure under per-visit billing, and if either site is outside the metro area it attracts loading on every single visit. A retainer's per-user pricing is indifferent to which building the user sits in. The exception is a genuinely tiny second site — a few users, no infrastructure — which is often best covered by dispatch or prepaid hours sitting on top of a retainer that covers the main office.
How do we forecast next year's spend under each?
Under a retainer, forecast headcount; the cost follows it almost linearly, which is why finance teams like it. Under per-visit, forecast events, and remember that event count does not track headcount — it tracks change. Office moves, new starters in batches, hardware refreshes and new sites all spike the physical layer without changing the user count. If your coming year contains any of those, model them explicitly rather than extrapolating last year's average. If you would like help putting your own numbers through this method, get in touch with twelve months of ticket history rather than a headcount.
Share:
Ready to take action?
Turn these insights into a roadmap for your business.
Book a 15-minute no-obligation consultation with our APAC IT experts. We'll review your current setup and provide a tailored IT roadmap within 24 hours.
Free Checklist
10 Critical Checks Before Expanding IT to Greater China
PIPL compliance, network segmentation, bilingual helpdesk setup, and more — everything your IT team needs before Day 1 in China.
Request the checklist →📬 Monthly Asia IT Insights
China compliance updates, cybersecurity alerts, and IT tips for APAC teams — once a month.
No spam. Unsubscribe anytime.
Related Articles
Aug 28, 2026
Too Small for a Full IT Contract, Too Big to Wing It: A Hong Kong Boutique Agency's Token Story
Jul 31, 2026
Onsite vs Remote IT Support in Hong Kong: Which Model Fits Your Office?
Jul 13, 2026
IT Support SLA Priority Levels Explained: P1, P2, P3 & P4 Response and Resolution Times