it-service@brocent.com
Availability Monitoring
Uptime, outages and capacity headroom across servers, business applications, storage, network devices and internet links for the August 2026 reporting period.
Trading-hours availability was unbroken; the internet link and NAS capacity are the risks
Every tier-1 service — market-data application, file server, internet access and identity — was available for the whole of Hong Kong trading hours across all 21 trading days. Overall availability across the 24 monitored items was 99.94%, against a 99.5% commitment. Three unplanned outages occurred, all outside trading hours, with a mean time to restore of 41 minutes.
Two capacity issues will become availability issues if left. The Hong Kong office has a single internet circuit with no failover, so any carrier fault stops market data, cloud access and voice together; and the NAS is at 87% capacity with roughly eleven weeks of headroom at the current growth rate. Both need a decision this quarter rather than a technical fix.
Availability position for the period
Availability is calculated on the monitored check, not on device power state: a server that responds to ping but whose application port is closed counts as unavailable for that service.
Four items require action
Evidence. A single 500 Mbps circuit terminates on the firewall HA pair. Peak utilisation is 62%, so bandwidth is not the issue — resilience is. The carrier reported two brief regional faults in the period; neither affected this building, which is luck rather than design.
Impact. A circuit fault during trading hours would simultaneously stop market data, Microsoft 365, cloud infrastructure access and hosted voice. There is no degraded mode: the office stops working. This is the single largest availability risk in the estate.
Recommendation. Add a second circuit from a different carrier on a different building entry, with automatic failover on the existing firewall pair. A 200 Mbps backup circuit is sufficient for degraded operation. BCS will provide two carrier quotations with the September report.
Evidence. The NAS is at 87% of 48 TB usable, growing 0.9 TB per month driven by research data. At that rate it passes the 95% threshold in the week of 2026-11-16, beyond which snapshot and rebuild behaviour degrades.
Impact. A full volume stops writes for the research team and causes backup jobs to fail on the same source, so the capacity issue becomes both an availability and a data-protection issue at once.
Recommendation. Either expand the array by one shelf or archive research data older than two years to cold cloud storage. BCS recommends the archive route: it is cheaper, reduces backup volume, and BCS will identify the candidate datasets for client approval.
Evidence. Shared documents are served by one physical server with no cluster partner and no warm standby. Recovery depends on a restore, which the Backup report verifies at 3 hours 12 minutes — below the four-hour objective but well beyond a trading-day tolerance.
Impact. A hardware failure at 09:00 would leave the firm without shared documents for most of a trading day, even though the recovery-time objective is technically met.
Recommendation. Continue the migration of active document sets to SharePoint, which removes the dependency without new hardware. Two of six shares have moved; propose the remaining four in the Q4 plan.
Evidence. One of the two Shenzhen access points logged four unexpected restarts, each under two minutes. Power-over-Ethernet delivery is at the edge of budget on the out-of-support access switch that feeds it.
Recommendation. Move the access point to a port on the supported switch; if restarts persist, replace the access point under warranty. Both are covered by the switch replacement already proposed.
The application server ran above 90% memory for four days in July, causing two market-data client disconnections. Memory was increased from 16 GB to 32 GB on 2026-08-03. Peak utilisation since is 54% and no disconnection has recurred.
Five of five service levels met
| Committed service level | Target | Actual | Status | Note |
|---|---|---|---|---|
| Tier-1 trading-hours availability | 100% | 100% | MET | 21 trading days, 4 services |
| Overall monthly availability | ≥ 99.5% | 99.94% | MET | 24 monitored items |
| Outage acknowledged | ≤ 15 minutes | 4 min worst | MET | 3 outages |
| Tier-1 restore time | ≤ 2 hours | 58 min worst | MET | File server outage 08-16 |
| Monthly report issued | By 5th working day | 3rd working day | MET | Issued 2026-09-03 |
Three-period movement
| Metric | Jun 2026 | Jul 2026 | Aug 2026 | Direction of travel |
|---|---|---|---|---|
| Overall availability | 99.81% | 99.88% | 99.94% | Improving; above commitment in all three periods. |
| Unplanned outages | 6 | 5 | 3 | Falling as the memory and firmware faults were fixed. |
| Mean time to restore | 72 min | 55 min | 41 min | Improving with runbooks for the top three fault types. |
| Trading-hours outages | 1 | 0 | 0 | Clean for two periods. |
| NAS capacity used | 79% | 83% | 87% | Rising 4 points per period — the binding constraint. |
| Circuit peak utilisation | 58% | 60% | 62% | Stable; bandwidth is not the concern, resilience is. |
Red above 85%, amber above 60%, green below. Only the NAS is in the red band; everything else has at least a year of headroom at current growth.
Availability by monitored item
Outage record
| Item | Start | Duration | Tier | Cause and resolution |
|---|---|---|---|---|
| File server | 2026-08-16 22:14 | 58 min | 1 | Storage controller firmware fault after a scheduled reboot; controller reseated and firmware rolled forward. Saturday night, no user impact. |
| Cloud VM 01 | 2026-08-21 03:40 | 37 min | 2 | Provider host maintenance without notice; VM restarted automatically on a healthy host. Reporting workload only. |
| Wireless AP (SZ) | 2026-08-27 11:02 | 28 min | 3 | Repeated restarts traced to PoE budget on the out-of-support switch; AP moved to a different port — see finding. |
Planned maintenance is excluded from availability when it falls in the agreed window and was notified 48 hours ahead. Two planned windows were used in the period, both on Sunday mornings, totalling 4 hours 20 minutes.
| Item | Class | Tier | Checks | Availability | Outages | Worst | Note |
|---|---|---|---|---|---|---|---|
| ACM-SRV-01 | Server | 1 | 6 | 99.87% | 1 | 58 min | Controller firmware |
| ACM-SRV-02 | Server | 1 | 7 | 100% | 0 | — | Memory upgraded 08-03 |
| Market-data app | Application | 1 | 5 | 100% | 0 | — | Synthetic login check |
| Accounting app | Application | 2 | 4 | 100% | 0 | — | Synthetic login check |
| Microsoft 365 | SaaS | 1 | 6 | 100% | 0 | — | Mail, Teams, SharePoint |
| ACM-NAS-01 | Storage | 2 | 8 | 100% | 0 | — | 87% capacity — finding |
| ACM-FW-01/02 | Firewall | 1 | 8 | 100% | 0 | — | HA pair; failover tested |
| HK internet circuit | WAN | 1 | 6 | 100% | 0 | — | No failover — finding |
| ACM-SW-C1/C2 | Core switch | 1 | 8 | 100% | 0 | — | Stacked pair |
| ACM-SW-A1/A2 | Access switch | 3 | 6 | 100% | 0 | — | A2 out of support |
| ACM-AP-01–04 | WIFI AP | 3 | 8 | 100% | 0 | — | HK office |
| ACM-AP-05/06 | WIFI AP | 3 | 4 | 99.94% | 1 | 28 min | AP-06 restarts — finding |
| Cloud VM 01 | Cloud server | 2 | 5 | 99.92% | 1 | 37 min | Provider maintenance |
| Cloud VM 02 | Cloud server | 2 | 5 | 100% | 0 | — | Reporting workload |
| Cloud database | Managed DB | 2 | 4 | 100% | 0 | — | Point-in-time recovery on |
| Site-to-site VPN | WAN | 2 | 6 | 100% | 0 | — | HK to SZ tunnel |
Sixteen rows cover the 24 monitored items; paired devices are shown as one row. Tier 1 items are those where an outage stops trading-hours work, tier 2 degrades it, tier 3 is localised.
Resilience position by service
| Service | Redundancy | Failover | Last tested | Position |
|---|---|---|---|---|
| Firewall / perimeter | HA pair | Automatic | 2026-08-10 | Failover tested during the maintenance window; 4-second interruption. |
| Core switching | Stacked pair | Automatic | 2026-06-14 | Adequate; next test scheduled with the access switch replacement. |
| Internet access | None | None | — | Single circuit — the estate's largest availability risk. |
| File services | None | Restore only | 2026-08-09 | 3 h 12 m restore verified; migration to SharePoint in progress. |
| Market-data application | Vendor cloud + local | Manual | 2026-07-19 | Traders can use the vendor web client if the local server fails. |
| Identity / Microsoft 365 | Provider | Automatic | — | Provider-managed; break-glass account held offline. |
| Voice | Provider + mobile | Manual | 2026-08-10 | Calls divert to mobiles; dependent on the same internet circuit. |
Open items carried between periods
| Ref | Risk | Severity | Owner | Due | Status | Next action |
|---|---|---|---|---|---|---|
| AVAIL-R01 | Single internet circuit — total office outage on any carrier fault | CRITICAL | Client COO | 2026-10-31 | Open | BCS to supply two carrier quotations for a diverse backup circuit. |
| AVAIL-R02 | NAS capacity exhaustion — approximately 11 weeks of headroom | HIGH | Client IT + BCS | 2026-10-15 | In progress | BCS to identify archive candidates; client to choose archive or expand. |
| AVAIL-R03 | File server single point of failure — recovery by restore only | MEDIUM | BCS + Client IT | 2026-12-31 | In progress | Migrate the remaining four shares to SharePoint. |
| AVAIL-R04 | Voice depends on the same circuit — no independent path | MEDIUM | Client COO | 2026-10-31 | Open | Resolved by the backup circuit; document the mobile divert meanwhile. |
| AVAIL-R05 | Out-of-support access switch — PoE instability and no firmware fixes | LOW | Client COO | 2026-11-30 | Open | Approve replacement — see CONF-R04. |
Next period commitments
| BCS Support Center will | Client is asked to |
|---|---|
| Provide two diverse-carrier quotations with failover design and cost. | Decide on the backup internet circuit by 2026-10-31. |
| Identify research datasets eligible for cold archive and size the saving. | Choose between NAS expansion and archiving older research data. |
| Move the Shenzhen access point to a supported switch port. | Approve the remaining SharePoint share migrations for Q4. |
| Report capacity headroom with weeks-to-threshold for every tier-1 item. | Confirm the trading-hours definition for the 2027 service schedule. |
How this report was produced
| Source | Extracted | Records | Used for |
|---|---|---|---|
| Monitoring platform | 2026-09-01 01:05 HKT | 96 checks | Uptime per check, outage start and end, acknowledgement times. |
| Synthetic application checks | 2026-09-01 01:05 HKT | 3 workflows | Application-level availability beyond host reachability. |
| SNMP performance history | 2026-09-01 01:20 HKT | 31 days | CPU, memory, interface and storage utilisation peaks. |
| Storage array telemetry | 2026-09-01 01:30 HKT | 1 array | Capacity consumed and growth rate for the headroom forecast. |
| Carrier portal | 2026-08-31 | 2 notices | Circuit utilisation and carrier fault notifications. |
| Incident tickets | 2026-08-31 | 3 outages | Cause, action taken and restoration evidence. |
Method and definitions
Availability is measured per monitored check at 60-second polling and aggregated by item, then by tier. A service is available only when its own check passes: host reachability alone is never treated as availability for an application.
Trading hours are 08:30 to 17:30 Hong Kong time on Hong Kong business days, 21 days in this period. Tier-1 availability is reported separately for that window because an outage at 03:00 and one at 10:00 are not equivalent events.
Capacity headroom is expressed as weeks until a defined threshold at the trailing three-month growth rate: 95% for storage, 80% sustained for CPU and memory, 70% peak for network interfaces. Forecasts are linear and deliberately conservative.
Exclusions. Planned maintenance inside the agreed window with 48 hours' notice is excluded. End-user devices are not in availability scope — they are covered in the IT Asset Management report. Third-party SaaS availability is reported as observed by our checks, not as claimed by the provider.
This copy is an anonymised sample prepared for illustration. The client name, device names, application names, carrier and provider identities have been replaced with fictitious or generic values; availability figures, ratios, dates and findings reflect a representative managed estate. No real client data appears in this document.
BCS Support Center · it-service@brocent.com
Questions clients ask about this report
Per monitored check at 60-second polling, aggregated by item and by service tier. A server that answers ping but whose application port is closed counts as unavailable — host reachability is never reported as application availability.
Because an outage at 03:00 and one at 10:00 are not the same event for a financial firm. Tier-1 availability is stated for Hong Kong trading hours as well as for the full month.
Yes — capacity headroom is expressed as weeks until a defined threshold at the trailing three-month growth rate, so a storage array filling up is raised as a scheduled decision rather than as an incident.
See what your own Availability Monitoring report would say
A free IT health check produces a first version of this report against your real estate, at no cost and with no obligation. It takes about a week and needs a few hours of your team’s time.