apatrickv.com/network STATUS: SOMEHOW OPERATIONAL

Network Stuff

A random collection of things that stopped working

What Is This?

This is not a networking tutorial.

It is a collection of real problems I've encountered while working with networks, telecom systems, customers, vendors, and the occasional piece of equipment that appears to have been designed by someone who has never had to troubleshoot it.

The cases are arranged roughly by OSI layer because it makes the page look organised. The actual troubleshooting process is usually less organised.

IMPORTANT: If the problem is not where you expect it to be, keep looking.

Layer 01 // Physical

The Rabbit Incident

L1

Back when we were still using point-to-point duplex fiber, I ran into a fault that stayed with me.

SNMP was showing the port down. The CPE LED, however, was showing broadband UP. So one side said there was no link while the other side apparently disagreed.

The optical readings made it stranger. We had about 8 dB in the direction of the CPE and essentially 0 dB in the direction of the switch.

Time for a light source and some actual testing.

The test showed a broken fiber in one of the two fibers making up the duplex pair. The break was about 1.5 metres behind the cupboard.

The culprit?

A rabbit.

It had apparently decided that the cable looked slightly more edible than it actually was.

INCIDENT REPORT
SNMP ................. PORT DOWN CPE LED .............. BROADBAND UP OPTICAL POWER ........ ASYMMETRIC LIGHT SOURCE TEST .... FIBER FAULT FAULT LOCATION ....... ~1.5 M BEHIND CUPBOARD ROOT CAUSE ............ RABBIT REPAIR ................ REPLACE HALF-FAULTY CABLE RABBIT ................. SPARED
LESSON: Start at Layer 1. The network does not care what the monitoring system thinks it should be doing.

Layer 02 // Data Link

The Friday Afternoon Incident

L2

Before we started using PON, our access network was mostly point-to-point. The competition had already moved to PON, so sometimes a customer leaving us meant a technician would simply move all the cables from the old CPE to the new one.

Sometimes that included the uplink cable to the media converter.

Which meant a customer's equipment could suddenly start talking directly to the network in ways it absolutely should not.

One of the more memorable symptoms was DHCP assignments escaping into places where DHCP assignments had no business being.

When I started working there, I encountered this within my first month. I checked whether DHCP Snooping was enabled.

It wasn't.

I quickly had a meeting with my boss and convinced him that we should implement DHCP Snooping and Option 82 properly.

Implementation started on a Friday at around 15:00.

There was one detail I had not yet discovered: two of the switches only supported 32 CPEs in DHCP snooping mode. They were both 24-port switches and were daisy-chained.

Even better, one of them was carrying a wireless point-to-multipoint network with roughly 50 clients.

About half an hour later, the phone started ringing.

Random customers could no longer get IP addresses.

What started as a sensible security improvement became an afternoon spent figuring out why perfectly innocent customers had suddenly stopped receiving DHCP.

FRIDAY CHANGE WINDOW
START ................ 15:00 OBJECTIVE ............. DHCP GUARDING OPTION 82 ............. NOT ENABLED SWITCH LIMIT .......... 32 CPE ACTUAL CLIENT COUNT ... ~50 CALLS .................. INCOMING TIME ................... FRIDAY AFTERNOON REGRET ................. IMMEDIATE
LESSON: Know the hardware limits before deploying the feature. Also, Friday afternoon is apparently a network maintenance concept with a physical form.

Layer 03 // Network

The Overlapping Subnet Incident

L3

We were deploying a new site. Everything was ready and my boss wanted it live that morning.

Everything, that is, except for one subnet I had forgotten to account for in IPAM.

The new site wanted 100.64.19.0/24.

Unfortunately, an existing allocation was 100.64.16.0/22.

These two networks overlap.

You can probably guess what happened next.

Calls from customers in 100.64.16.0/22 started arriving, reporting that something weird was going on.

The actual problem was not particularly exotic. The exotic part was the fact that I had managed to schedule the problem into production before noticing it in IPAM.

IPAM INCIDENT
NEW SITE .............. 100.64.19.0/24 EXISTING NETWORK ....... 100.64.16.0/22 OVERLAP ................ YES CGNAT .................. INVOLVED CUSTOMER CALLS ......... YES IPAM CHECK ............. TOO LATE ROOT CAUSE ............. HUMAN
LESSON: Plan deployments in IPAM, not in the minds of sleep-deprived engineers first thing in the morning.

Layer 04 // Transport

The Two ASNs Incident

L4-ish

We had a transit provider that had merged with another ISP. From a RIPE perspective they were still operating under two separate ASNs.

Because our transit came from one ASN, we still wanted peering with the other.

They agreed and provided peering subnets.

We set the BGP preference so traffic would prefer the peering site where we had established the session.

The problem was that, functionally, the two networks were now one network hiding under two ASNs.

The result was an asymmetric path.

Upload traffic came through Site A.

Download traffic came through Site B.

Everything looked reasonable until you looked at the actual path the traffic was taking.

The issue was never properly fixed. To this day, we are limited to using their transit.

They don't provide a discount for the privilege.

PATH ANALYSIS
PROVIDER .............. ONE COMPANY ASN ................... TWO PEERING ............... SITE A UPLOAD ................ SITE A DOWNLOAD .............. SITE B SYMMETRIC ............. NO FIX ................... NONE DETECTED DISCOUNT .............. ALSO NONE DETECTED
LESSON: Routing policy describes paths. It does not guarantee that the network on the other side of the router has the topology you imagined.

Layer 05 // Session

No Evidence Yet

L5

I don't currently have a sufficiently entertaining real-world Layer 5 incident to put here.

Rather than invent one, this box is intentionally left as a placeholder.

DATABASE STATUS
REAL INCIDENT ........ NOT FOUND FICTION ............... DISABLED MEMORY SEARCH ......... RUNNING
LESSON: Knowing that you don't have the answer is also useful troubleshooting information.

Layer 06 // Presentation

Pending Investigation

L6

Layer 6 is currently awaiting a sufficiently ridiculous encoding, formatting, encryption, or representation failure from the field.

INCIDENT DATABASE
SUITABLE CASE ........ NOT LOCATED PLACEHOLDER .......... ACCEPTED OVERENGINEERING ....... POSSIBLE
LESSON: Do not manufacture evidence to make a page look complete.

Layer 07 // Application

The Shiny Buttons Incident

L7

A client decided to purchase VoIP from us.

I sent him the usual SIP information: username, password, SIP server and DID.

NORMAL SIP ENDPOINT PROVISIONING
TIME .................. ~30 SECONDS VENDOR ................ IRRELEVANT USERNAME .............. PROVIDED PASSWORD .............. PROVIDED SIP SERVER ............ PROVIDED DID ................... PROVIDED EXPECTED RESULT ....... DONE

He replied asking if I knew how to put it into UniFi Talk.

I found the guide and sent it to him.

A few days later I received a screenshot containing a collection of information that was technically related to SIP but somehow managed to communicate absolutely nothing useful.

I checked the PBX. The endpoint wasn't registered.

A few days passed. Eventually he asked me to come over and configure it with him, because apparently giving me access and letting me configure it remotely was too easy.

ROUND ONE

I entered the SIP credentials according to the documentation. After some clicking of shiny buttons, outbound calls worked.

Inbound calls did not.

I clicked around some more.

Nothing useful happened.

So I went back to the office and ran sngrep while placing an inbound call.

The SIP traffic showed something interesting: the From field was not being populated correctly.

I checked the documentation again and sent another email:

Set this field to test, try again, and let me know.

Three days later:

Can you come over and do it?

ROUND TWO

Back on site. Enter test. Make a test call.

Working.

Obviously, this was the end of the story.

IT WAS NOT THE END OF THE STORY

A few months later the client had updated, reset, or otherwise interacted with the device in a way that caused the configuration to stop working again.

I restored the configuration as closely as possible, but the interface had changed. Fields had moved. Options had been renamed.

Apparently the configuration interface had evolved.

Inbound calls worked.

Outbound calls did not.

The client eventually fell asleep on his couch while I was trying to figure it out.

Fair enough.

ROUND THREE

Back to the office.

sngrep showed that the field previously called From had apparently become Use username in Contact URI.

I changed it and tested again.

Inbound calls were still failing.

So I dropped another layer.

tcpdump showed the connection being dropped at the gateway.

Finally I checked the UniFi firewall rules.

An automatically generated rule was wrong.

Fix the rule.

Test.

Calls working.

FINAL DIAGNOSIS
PBX .................. ENDPOINT NOT REGISTERED SIP .................. INCORRECT HEADER UI ................... FIELDS CHANGED SNGRЕP ............... EVIDENCE FOUND TCPDUMP .............. CONNECTION DROPPED FIREWALL ............. AUTOGENERATED RULE WRONG FINAL RESULT .......... CALLS WORKING CLIENT ASLEEP ......... YES
LESSON: Don't troubleshoot the product. Troubleshoot the traffic. Documentation can tell you what a button is supposed to do. The packet doesn't care.

General Troubleshooting Protocol

01Believe the symptom, not the assumption.
02Start at Layer 1.
03Change one thing at a time.
04When clicking stops producing information, capture packets.
05Check the documentation.
06Check the actual traffic anyway.
07If the problem makes no sense, look for the thing nobody thought could possibly be wrong.
← [ HOME ]