[AusNOG] Telstra outage?

Mark Smith markzzzsmith at gmail.com
Tue Jul 14 19:02:31 AEST 2026


On Tue, 14 Jul 2026 at 18:03, Robert Hudson <hudrob at gmail.com> wrote:

> I'd suggest VMs are fine as a time source as long as they also sync from
> something else AND you're aware of how your hypervisor impacts time keeping
> and allow for that.
>
>
How do you allow for that?

Is accurate time a nice to have or a need to have?

I've seen a pretty large customer facing network access device reboot
periodically (avoiding naming the vendor because I don't think it was their
fault), e.g, about once every 6 weeks, causing many customer service
outages.

The only thing in the error logs before the reboot were messages about the
main controller card and the linecards' clocks getting out of sync with NTP
because upstream NTP wasn't reliable - switching stratums, time between the
two upstream NTP servers being significantly different (so they weren't
sync'd with each other either).

Switching off NTP on the device stopped it from rebooting. That's all we
could do because we couldn't get the upstream NTP service fixed because
they didn't care about it that much.

I figure it was because the messages between the main control card and
linecards are time stamped so they're processed in order (think about
deleting an old version of a route in a linecard before installing a new
one, rather than installing a new version of the route and then deleting
it), and if the clocks on the main controller card and line cards get too
out of whack, the device panic reboots to try to fix the issue.

So it seems a "nice to have" upstream NTP service that wasn't accurate and
people didn't really care about became a need to be accurate on a network
access device.

Regards,
Mark.



On Tue, 14 July 2026, 7:56 am Mark Smith, <markzzzsmith at gmail.com> wrote:
>
>> Going from the news reports they only had two.
>>
>> Two services ir devices for redundancy is usually enough.
>>
>> That's what I thought about NTP, until I read the following BCP.
>>
>> The ideal minimum number of NTP time sources for a client is *5*, due to
>> the way time sources are compared to try to identify the best and most
>> accurate one to follow, and then to have redundancy, which comes with
>> having an odd number of NTP sources.
>>
>> RFC8633, "Network Time Protocol Best Current Practices"
>> https://datatracker.ietf.org/doc/html/rfc8633
>>
>>
>> Also, don't use VMs as NTP time servers. The underlying virtualisation
>> layer can cause the VM's virtual timer interrupt to drift about.
>>
>> You want a hardware timer interrupt to be driving the OS's clock.
>>
>> If Telstra had had 5 Stratacom NTP servers, even if they were all old and
>> had the 20 years ago time fault, it's possible that due to different power
>> on times, the first one that failed and turned time back 20 years would
>> have been ignored by the others and clients because it became a bad time
>> source, and they may have been lucky enough to have time to mitigate the
>> remaining ones failing before it had an impact.
>>
>>
>> Regards,
>> Mark.
>>
>> On Tue, 14 July 2026, 15:30 Luke Thompson, <luke.t at tnc.works> wrote:
>>
>>> "Telstra had a nation-wide network outage last week that affected
>>> emergency services.
>>>
>>> The outage has been pinned on an obsolete Symmetricom SyncServer S300
>>> node, which manages time on the network but resets its 10-bit week counter
>>> to zero every 1024 weeks (just under 20 years) in a common and well
>>> understood GPS rollover bug that caused the device to reset to 2006.
>>>
>>> The SyncServer S300 was discontinued in 2016."
>>>
>>> Cheers,
>>>
>>> Luke Thompson, CTO
>>> The Network Crew P/L
>>>
>>> E: luke.t at tnc.works
>>> https://tnc.works
>>>
>>> On 9 July 2026 7:21:52 am Greg Price <greg at lakemountain.com> wrote:
>>>
>>>> The stuff I've read and heard is that it was a software update across
>>>> multiple systems that provide the reference sources from GNSS that went
>>>> wrong causing the the reference date/time to jump back decades, and the
>>>> wheels fall off.  This is from the patchwork of sources, but I'd like to
>>>> hear it definitively from engineers.
>>>>
>>>> On 9 Jul 2026, at 7:12 am, Phillip Grasso <phillip.grasso at gmail.com>
>>>> wrote:
>>>>
>>>> Do you think it was related to software patching breaking time
>>>> services? (NTP?)
>>>>
>>>> On Wed, 8 July 2026, 2:48 pm DaZZa, <dazzagibbs at gmail.com> wrote:
>>>>
>>>>> To paraphrase an oft-repeated meme
>>>>>
>>>>> It's not NTP
>>>>> There's no way it's NTP
>>>>> It was NTP
>>>>>
>>>>> Bets on this just being a speculator to take the heat off until they
>>>>> find out what really went wrong?
>>>>>
>>>>> D
>>>>>
>>>>> On Wed, 8 Jul 2026 at 14:34, Luke Thompson <luke.t at tnc.works> wrote:
>>>>> >
>>>>> > Telstra's CFO has pointed to time / NTP:
>>>>> >
>>>>> >
>>>>> https://www.theguardian.com/australia-news/2026/jul/08/telstra-outage-down-network-mobile-outages-today
>>>>> >
>>>>> > Seems just over 4 hours of impact, and so far they're unsure what
>>>>> caused it.
>>>>> >
>>>>> > Cheers,
>>>>> >
>>>>> > Luke Thompson, CTO
>>>>> > The Network Crew P/L
>>>>> >
>>>>> > E: luke.t at tnc.works
>>>>> > https://tnc.works
>>>>> >
>>>>> > On 8 July 2026 8:43:01 am Luke Thompson <luke.t at tnc.works> wrote:
>>>>> >
>>>>> >> & likewise, service restored at the moment.
>>>>> >>
>>>>> >> Cheers,
>>>>> >>
>>>>> >> Luke Thompson, CTO
>>>>> >> The Network Crew P/L
>>>>> >>
>>>>> >> luke.t at tnc.works
>>>>> >> https://tnc.works
>>>>> >>
>>>>> >>
>>>>> >> On 8/7/2026 8:40 am, lauricat at fastmail.fm wrote:
>>>>> >>>
>>>>> >>> Good morning
>>>>> >>>
>>>>> >>> Back here. Regional Vic  Boost and T prepaid mobile data.
>>>>> >>>
>>>>> >>> ----- Original message -----
>>>>> >>> From: lauricat at fastmail.fm
>>>>> >>> To: ausnog at lists.ausnog.net
>>>>> >>> Subject: [AusNOG] Telstra outage?
>>>>> >>> Date: Wednesday, 8 July 2026 7:44 AM
>>>>> >>>
>>>>> >>> G'day.
>>>>> >>>
>>>>> >>> How is T where you are?
>>>>> >>>
>>>>> >>> Regional Vic Boost network no calls other than emergency can be
>>>>> made.
>>>>> >>>
>>>>> >>> Cheers
>>>>> >>> Laurie
>>>>> >>>
>>>>> >> _______________________________________________
>>>>> >> AusNOG mailing list
>>>>> >> AusNOG at lists.ausnog.net
>>>>> >> https://lists.ausnog.net/mailman/listinfo/ausnog
>>>>> >
>>>>> >
>>>>> > _______________________________________________
>>>>> > AusNOG mailing list
>>>>> > AusNOG at lists.ausnog.net
>>>>> > https://lists.ausnog.net/mailman/listinfo/ausnog
>>>>> _______________________________________________
>>>>> AusNOG mailing list
>>>>> AusNOG at lists.ausnog.net
>>>>> https://lists.ausnog.net/mailman/listinfo/ausnog
>>>>>
>>>> _______________________________________________
>>>> AusNOG mailing list
>>>> AusNOG at lists.ausnog.net
>>>> https://lists.ausnog.net/mailman/listinfo/ausnog
>>>>
>>>>
>>>>
>>> _______________________________________________
>>> AusNOG mailing list
>>> AusNOG at lists.ausnog.net
>>> https://lists.ausnog.net/mailman/listinfo/ausnog
>>>
>> _______________________________________________
>> AusNOG mailing list
>> AusNOG at lists.ausnog.net
>> https://lists.ausnog.net/mailman/listinfo/ausnog
>>
>
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <https://lists.ausnog.net/pipermail/ausnog/attachments/20260714/84d9ad8f/attachment-0001.htm>


More information about the AusNOG mailing list