Network Stuff
A random collection of things that stopped working
What Is This?
This is not a networking tutorial.
It is a collection of real problems I've encountered while working with networks, telecom systems, customers, vendors, and the occasional piece of equipment that appears to have been designed by someone who has never had to troubleshoot it.
The cases are arranged roughly by OSI layer because it makes the page look organised. The actual troubleshooting process is usually less organised.
Layer 01 // Physical
The Rabbit Incident
L1Back when we were still using point-to-point duplex fiber, I ran into a fault that stayed with me.
SNMP was showing the port down. The CPE LED, however, was showing broadband UP. So one side said there was no link while the other side apparently disagreed.
The optical readings made it stranger. We had about 8 dB in the direction of the CPE and essentially 0 dB in the direction of the switch.
Time for a light source and some actual testing.
The test showed a broken fiber in one of the two fibers making up the duplex pair. The break was about 1.5 metres behind the cupboard.
The culprit?
A rabbit.
It had apparently decided that the cable looked slightly more edible than it actually was.
SNMP ................. PORT DOWN CPE LED .............. BROADBAND UP OPTICAL POWER ........ ASYMMETRIC LIGHT SOURCE TEST .... FIBER FAULT FAULT LOCATION ....... ~1.5 M BEHIND CUPBOARD ROOT CAUSE ............ RABBIT REPAIR ................ REPLACE HALF-FAULTY CABLE RABBIT ................. SPARED
Layer 02 // Data Link
The Friday Afternoon Incident
L2Before we started using PON, our access network was mostly point-to-point. The competition had already moved to PON, so sometimes a customer leaving us meant a technician would simply move all the cables from the old CPE to the new one.
Sometimes that included the uplink cable to the media converter.
Which meant a customer's equipment could suddenly start talking directly to the network in ways it absolutely should not.
One of the more memorable symptoms was DHCP assignments escaping into places where DHCP assignments had no business being.
When I started working there, I encountered this within my first month. I checked whether DHCP Snooping was enabled.
It wasn't.
I quickly had a meeting with my boss and convinced him that we should implement DHCP Snooping and Option 82 properly.
Implementation started on a Friday at around 15:00.
There was one detail I had not yet discovered: two of the switches only supported 32 CPEs in DHCP snooping mode. They were both 24-port switches and were daisy-chained.
Even better, one of them was carrying a wireless point-to-multipoint network with roughly 50 clients.
About half an hour later, the phone started ringing.
Random customers could no longer get IP addresses.
What started as a sensible security improvement became an afternoon spent figuring out why perfectly innocent customers had suddenly stopped receiving DHCP.
START ................ 15:00 OBJECTIVE ............. DHCP GUARDING OPTION 82 ............. NOT ENABLED SWITCH LIMIT .......... 32 CPE ACTUAL CLIENT COUNT ... ~50 CALLS .................. INCOMING TIME ................... FRIDAY AFTERNOON REGRET ................. IMMEDIATE
Layer 03 // Network
The Overlapping Subnet Incident
L3We were deploying a new site. Everything was ready and my boss wanted it live that morning.
Everything, that is, except for one subnet I had forgotten to account for in IPAM.
The new site wanted 100.64.19.0/24.
Unfortunately, an existing allocation was 100.64.16.0/22.
These two networks overlap.
You can probably guess what happened next.
Calls from customers in 100.64.16.0/22 started arriving, reporting that something weird was going on.
The actual problem was not particularly exotic. The exotic part was the fact that I had managed to schedule the problem into production before noticing it in IPAM.
NEW SITE .............. 100.64.19.0/24 EXISTING NETWORK ....... 100.64.16.0/22 OVERLAP ................ YES CGNAT .................. INVOLVED CUSTOMER CALLS ......... YES IPAM CHECK ............. TOO LATE ROOT CAUSE ............. HUMAN
Layer 04 // Transport
The Two ASNs Incident
L4-ishWe had a transit provider that had merged with another ISP. From a RIPE perspective they were still operating under two separate ASNs.
Because our transit came from one ASN, we still wanted peering with the other.
They agreed and provided peering subnets.
We set the BGP preference so traffic would prefer the peering site where we had established the session.
The problem was that, functionally, the two networks were now one network hiding under two ASNs.
The result was an asymmetric path.
Upload traffic came through Site A.
Download traffic came through Site B.
Everything looked reasonable until you looked at the actual path the traffic was taking.
The issue was never properly fixed. To this day, we are limited to using their transit.
They don't provide a discount for the privilege.
PROVIDER .............. ONE COMPANY ASN ................... TWO PEERING ............... SITE A UPLOAD ................ SITE A DOWNLOAD .............. SITE B SYMMETRIC ............. NO FIX ................... NONE DETECTED DISCOUNT .............. ALSO NONE DETECTED
Layer 05 // Session
No Evidence Yet
L5I don't currently have a sufficiently entertaining real-world Layer 5 incident to put here.
Rather than invent one, this box is intentionally left as a placeholder.
REAL INCIDENT ........ NOT FOUND FICTION ............... DISABLED MEMORY SEARCH ......... RUNNING
Layer 06 // Presentation
Pending Investigation
L6Layer 6 is currently awaiting a sufficiently ridiculous encoding, formatting, encryption, or representation failure from the field.
SUITABLE CASE ........ NOT LOCATED PLACEHOLDER .......... ACCEPTED OVERENGINEERING ....... POSSIBLE
Layer 07 // Application
The Shiny Buttons Incident
L7A client decided to purchase VoIP from us.
I sent him the usual SIP information: username, password, SIP server and DID.
TIME .................. ~30 SECONDS VENDOR ................ IRRELEVANT USERNAME .............. PROVIDED PASSWORD .............. PROVIDED SIP SERVER ............ PROVIDED DID ................... PROVIDED EXPECTED RESULT ....... DONE
He replied asking if I knew how to put it into UniFi Talk.
I found the guide and sent it to him.
A few days later I received a screenshot containing a collection of information that was technically related to SIP but somehow managed to communicate absolutely nothing useful.
I checked the PBX. The endpoint wasn't registered.
A few days passed. Eventually he asked me to come over and configure it with him, because apparently giving me access and letting me configure it remotely was too easy.
ROUND ONE
I entered the SIP credentials according to the documentation. After some clicking of shiny buttons, outbound calls worked.
Inbound calls did not.
I clicked around some more.
Nothing useful happened.
So I went back to the office and ran sngrep while placing an inbound call.
The SIP traffic showed something interesting: the From field was not being populated correctly.
I checked the documentation again and sent another email:
test, try again, and let me know.Three days later:
ROUND TWO
Back on site. Enter test. Make a test call.
Working.
Obviously, this was the end of the story.
IT WAS NOT THE END OF THE STORY
A few months later the client had updated, reset, or otherwise interacted with the device in a way that caused the configuration to stop working again.
I restored the configuration as closely as possible, but the interface had changed. Fields had moved. Options had been renamed.
Apparently the configuration interface had evolved.
Inbound calls worked.
Outbound calls did not.
The client eventually fell asleep on his couch while I was trying to figure it out.
Fair enough.
ROUND THREE
Back to the office.
sngrep showed that the field previously called From had apparently become Use username in Contact URI.
I changed it and tested again.
Inbound calls were still failing.
So I dropped another layer.
tcpdump showed the connection being dropped at the gateway.
Finally I checked the UniFi firewall rules.
An automatically generated rule was wrong.
Fix the rule.
Test.
Calls working.
PBX .................. ENDPOINT NOT REGISTERED SIP .................. INCORRECT HEADER UI ................... FIELDS CHANGED SNGRЕP ............... EVIDENCE FOUND TCPDUMP .............. CONNECTION DROPPED FIREWALL ............. AUTOGENERATED RULE WRONG FINAL RESULT .......... CALLS WORKING CLIENT ASLEEP ......... YES