What we've learned from the external investigation of our July outage
On 8 July, Telstra experienced a mobile network outage that disrupted some voice, messaging and data services for customers across Australia, including impacts to some Triple Zero calls.
It was an outage that should not have happened.
We let our customers down and we fell short of what Australians rightly expect from us.
In the days after the outage, I committed to being transparent about what happened, what we'd learn from it, and what we'd do differently.
Today, we're releasing the findings of an external expert investigation conducted by Technology Audit Partners (TAP).
What the investigation found
The investigation confirmed the outage was triggered by a specific technical event, consistent with what we said at the time: incorrect date information propagated through parts of our mobile network following planned maintenance on the network timing system.
Most significantly, TAP found the outage was primarily a result of us not treating network timing as a critical capability within the network (or a ‘sovereign function’) requiring the highest levels of oversight and protection1.
As it relates to the network timing system, TAP identified that:
Ownership for this function was not sufficiently defined, which led to weakness in this area and not treating it as a high-risk function of the network for the purposes of design, operations and resource prioritisation.
The architecture supporting network timing had evolved over time without adequate end-to-end oversight.
There was insufficient technical expertise for network timing systems and the need for stronger competency and more curiosity to investigate when something looked wrong.
There were shortcomings in process discipline including change management, configuration control, incident management, documentation and follow through.
The investigation also found that once the major incident response was initiated, Telstra's teams managed a highly complex recovery effectively. However, gaps in ownership, visibility and operational support of the network timing system delayed the identification of the problem in the early stages of the outage.
What we've already done
We didn't wait for TAP’s findings to begin making changes.
Since July, we’ve taken a series of immediate actions to strengthen the resilience of the timing system in our mobile network.
We have migrated services away from the previous network time protocol (NTP) servers to our strategic system across all three sites. We’ve also added additional monitoring and alarm capabilities in our network, introduced additional testing of network changes in lab environments, and worked closely with our vendors to continue uplifting our operational and change processes.
At the same time, we established a company-wide program overseeing remediation, network resilience improvements and implementation of the findings arising from the outage.
What happens next
We accept the findings of TAP’s investigation.
Many of the findings align with issues we identified through our remediation work. Others provide valuable external insights into where we need to strengthen our network and management processes.
Our focus is now on implementation.
That includes:
Completing a review of critical non-timing related functions across our network to ensure the right level of priority has been applied
Reviewing our vendor notification management processes to ensure that any future alerts are given appropriate focus
Assessing our platform alarm instructions and service assurance arrangements to ensure our technical teams can action all issues without delay
We will also continue to work with customers on any complaints or claims associated with the outage.
Our responsibility
Modern networks are complex, but complexity is not an excuse.
Our customers expect us to operate reliable and resilient networks. They expect us to learn when things go wrong. And they expect us to be transparent about both.
This outage should not have happened, and the findings released today help explain why it did.
Our commitment is to use what we have learned from this outage not only to address what went wrong, but to make Telstra stronger and our services even more resilient and reliable for our customers.
You can read TAP’s findings attached below.
1 In TAP’s findings, a ‘Sovereign Function’ means a network function so critical that, if it fails, it could cause a network-wide outage. This is different from the more common use of ‘sovereignty’ in telecommunications, which usually refers to network ownership, control or where data is held.