BGP: The Protocol That Runs the Internet (Whether You Like It or Not)

Pour something good. This one’s worth it.

Border Gateway Protocol, or BGP, has been running the internet since 1989. Your ISP uses it. Amazon uses it. Google uses it. The backbone of every packet that leaves your network and crosses into someone else’s depends on it. And there’s a decent chance you’ve been avoiding it because someone told you early in your career that BGP was “a service provider thing.”

They were wrong. And this series is going to prove it.

What BGP Actually Is

BGP stands for Border Gateway Protocol. It’s the protocol that handles routing between autonomous systems, which is a fancy way of saying it manages how networks talk to other networks that they don’t own or control.

If you’ve spent most of your time with OSPF or EIGRP, you’re used to link-state or distance-vector protocols. BGP is neither. It’s a path-vector protocol, and that distinction matters more than most people realize.

Where OSPF builds a map of the network and calculates shortest paths, BGP builds a table of paths and applies policy to pick the best one. BGP doesn’t care about bandwidth or delay, by default. It cares about which autonomous systems a packet has to pass through to reach a destination, and it applies rules you define to decide which path wins.

That policy control is the whole point of BGP. It’s also why it’s more complex than anything else in your routing stack.

A Quick History Because It Actually Matters

The internet’s first routing protocol for connecting separate networks was EGP, Exterior Gateway Protocol. It worked fine when the internet was small. Then the internet stopped being small.

BGP-4 first appeared in RFC 1654 in 1994, refined by RFC 1771 in 1995, and eventually standardized in RFC 4271. but the core mechanics haven’t changed dramatically since then. Classless routing, route aggregation, and policy control were all baked in from the beginning.

The fact that a protocol designed in the early 90s is still holding the global routing table together tells you two things. First, the design was solid. Second, the bar for replacing it is so high that nobody’s come close.

Where You’ll Actually See BGP

Here’s where people get surprised. BGP isn’t just for ISPs and carriers. You’re going to run into it in places that have nothing to do with service provider networks.

Multi-homed enterprises. If your company has two ISPs and you want actual control over how traffic goes out and comes in, you need BGP. Static routes and hope is not a traffic engineering strategy.

Data center fabrics. Spine-leaf architectures running BGP as the underlay protocol have become standard. Microsoft, Google, Facebook, and pretty much every large data center operator runs BGP as the IGP replacement in their fabric. RFC 7938 made this an actual recommendation. If you’re touching data center networking, you will touch BGP.

SD-WAN handoffs. Most enterprise SD-WAN solutions are peering with BGP at the edge to pull in routes from the WAN fabric. The box might abstract it away, but BGP is underneath.

Cloud connectivity. AWS Direct Connect, Azure ExpressRoute, Google Cloud Interconnect. All of them use BGP sessions to exchange routes between your on-prem environment and the cloud. You can get away with letting the cloud provider manage it, but understanding what’s happening is the difference between troubleshooting in 10 minutes and staring at the console for two hours.

What Problem BGP Actually Solves

Strip away all the complexity for a second and look at the core problem.

You have a network. Someone else has a network. You want to talk to each other. And you both want control over how that traffic flows. You don’t want your neighbor deciding that all your traffic should route through some third network you’ve never heard of. You want to apply your own policies.

BGP gives you that control. When you receive a BGP update, you decide what to accept. When you send one, you decide what to advertise. Inbound traffic can be influenced with AS path prepending. Outbound traffic can be influenced with local preference. You can filter prefixes you don’t want to accept. You can summarize routes you don’t want to advertise in detail.

OSPF can’t do any of that. EIGRP can’t do that. Static routes can sort of do it, but not at scale. BGP was built for exactly this problem.

The Honest Disadvantage

BGP will absolutely let you configure something that works perfectly in lab and destroys production. And it will do it quietly. There’s no immediate error. The session stays up. The routes look fine. And then three hours later someone figures out that half your traffic is taking a 14-hop path through Eastern Europe because of one wrong route-map entry.

The complexity is real. The debugging takes practice. A configuration change that pulled Facebook’s routes off the internet caused the Facebook outage in October 2021 that took down Instagram, WhatsApp, and Oculus for six hours. A BGP route leak in 2019 caused 70,000 routes from Cloudflare and others to detour through a small Pennsylvania ISP and take down large portions of the internet for an hour and a half.

These weren’t bugs in BGP. They were operator errors. BGP did exactly what it was configured to do. The configurations were just wrong.

That’s the thing about BGP. It’s not forgiving. It expects you to know what you’re doing. This series exists because more engineers should.

What’s Coming

Over the next several weeks we’re going through the whole stack. eBGP versus iBGP. Autonomous system numbers. Path selection. Route filtering. Route reflectors. Confederations. Address families.

Each one builds on the last. By the end you’ll have a real mental model for how BGP works, why it behaves the way it does, and how to configure it without breaking things.

Leave a Comment

Your email address will not be published. Required fields are marked *