Skip to content
Protocolzone Protocolzone

Service line

24×7 Operations Desk

Monitoring, incident response & release support

Discuss your project →

Wagering, payments and data platforms fail on the schedule that hurts most: overnight, on a public holiday, in the last minutes before a Saturday metro meeting. The question is not whether someone is monitoring. It is whether the person who answers can read the code, has the access to act, and knows what the system is supposed to be doing.

The Operations Desk is staffed by our own engineers. It is not a subcontracted first line reading a script and creating a ticket for somebody else.

What the desk covers

We run the platforms we build, and we take on platforms we did not build. The second case is the harder one and we treat it as an engagement in its own right: read the system, find out what is actually monitored, write the runbooks that were never written, and produce an honest list of what will break before it does.

Cloud-native, hybrid and legacy systems are all in scope. Legacy usually means undocumented rather than old, and documenting it is part of the job.

Coverage

The desk runs continuously. Cover spans Australian racing hours, including Saturday metro meetings and overnight processing windows, and shift handovers are written rather than verbal.

Response targets are set per engagement against your severity definitions. We would rather agree them with you than publish a number here that turns out not to apply to your system.

How we work inside your process

The desk fits your existing architecture, ticketing and change governance rather than requiring a parallel one. Escalation paths, approval requirements and out-of-hours authority are agreed during onboarding. During an incident, nobody should be working out who is allowed to authorise a rollback.

Incidents produce changes

Every incident gets a record, a severity and an owner. Recurring incidents are defects, and they go into the backlog of whoever owns that service with the evidence attached. An operations desk that absorbs the same failure every fortnight without fixing it is a cost, not a service.

When this is the wrong call

If your system has no monitoring and no runbooks, buying overnight cover first is the wrong order of work. The desk cannot see what is not instrumented. In that case the engagement starts with an instrumentation and runbook piece, usually alongside Platform to Platform, and cover starts once there is something to watch.

Capabilities

What sits inside this service line.

Monitoring and alerting

Alerts written against business outcomes rather than machine metrics: a feed that has stopped moving, a settlement queue that is growing, a partner interface returning errors. Thresholds are tuned so a page means something.

Incident response

A named engineer on shift, a documented triage path, and a running incident record. Severity decides who gets woken up, and that mapping is agreed with you before it is needed rather than negotiated during an outage.

Release and change support

Deployments, migrations and configuration changes executed inside your change process, with rollback tested rather than assumed.

Customer-owned platforms

We take on systems we did not build: cloud-native, hybrid or legacy. Onboarding starts with reading the code and writing the runbooks that do not exist yet.

Shift handover

Written handover between shifts covering open incidents, watch items and anything deferred, so the next engineer starts informed instead of re-diagnosing.

Continuous improvement

Recurring incidents are treated as defects with an owner, not as operational weather. Post-incident reviews produce a code or configuration change, or they produce nothing.

Deliverables

What you actually receive.

Artefacts, not adjectives. Every item on this list is something you can point at when the engagement ends.

  • Named shift coverage with an agreed escalation matrix
  • Runbook set per service, maintained as the system changes
  • Monitoring and alerting configuration tuned against business signals
  • Incident records, severity classification and post-incident reviews
  • Written shift handover log
  • Change and release execution inside your governance process
  • Onboarding assessment and gap list for platforms we did not build

Read next

24×7 Operations Desk in practice.

Use cases

Case studies

The other two lines

Plan your next platform or improvement.

Tell us what you want to build, connect or improve. We can start with a business goal or work through the technical constraints you already know.