📑 Daftar Isi
Massive DDoS Attack: International Link Drop, Rack Switch Failure — DRC Rescue Story
It was 11 PM on a Saturday. The quietest shift of the week. Clients were off enjoying their weekends, support tickets were at an all-time low, and I was sipping coffee while half-watching Netflix on my second monitor. Then — BEEEP BEEEP BEEEP. Grafana alerts started screaming. My screen went from green to red in seconds. International inbound traffic graph went vertical — 5 Gbps, 10 Gbps, 20 Gbps, 40 Gbps — hitting the link capacity hard. “What the…” I muttered. New game launch? No. This was a massive DDoS attack — the kind you read about in post-mortems but never expect to deal with firsthand.
Who was targeting us? No idea at that moment. What I knew was this: five minutes earlier, international traffic was a peaceful 2-3 Gbps. Now it was pegged at 40 Gbps and still climbing. Packet loss jumped from 0.01% to 35% in minutes. Ping to our Singapore node — normally 35ms — spiked to 450ms. Some requests timed out entirely. API calls from overseas clients started failing left and right. Our NOC Telegram group exploded: “Hey, website is down from Europe!” “VPN connection to office dropped!” “Email queue is piling up, what’s going on?” Panic? Definitely. But here’s the thing — panic doesn’t fix anything. Procedure does. And we had a Disaster Recovery Center (DRC) protocol ready to roll.
Within minutes, our Grafana dashboard turned into a war room. Traffic analysis from the netflow data that managed to get captured before the tools got overloaded showed a pattern: UDP flood mixed with SYN flood from thousands of different source IPs. This was a distributed reflection attack (DRDoS) with massive amplification factor — likely using NTP or CLDAP protocols where the bandwidth multiplier can reach 10-50x. What that means in plain English: the attacker only needed a small amount of bandwidth to generate a tsunami of traffic toward the target. And the target was clear — our main upstream provider. The immediate impact was brutal: international link saturated completely, BGP sessions started flapping, routes to overseas destinations were getting withdrawn. Clients from Europe, America, Australia started calling — their websites were completely inaccessible. For businesses that rely on online revenue, every minute of downtime meant real money lost.
But the night wasn’t done with us yet. One of our Site Leads called from DC-2 — a switch rack was showing strange symptoms. Temperature sensor on the main switch had spiked from a normal 37°C to 58°C in just 30 minutes. Two out of four fan trays on that switch had failed — probably degrading for a while, and the extra traffic load from the DDoS was the final straw. Spanning Tree Protocol started sending Topology Change Notifications continuously. Several servers in that rack started disconnecting and reconnecting in 5-minute cycles. This was a perfect storm — external attack combined with internal hardware failure. The NOC team had to split focus: one group handling DDoS mitigation, another handling the switch rack emergency. We set up a dedicated incident response channel — no mixing with general chat to avoid chaos and keep a clean history of every action taken.
Two major problems in one night — a massive DDoS attack hitting our upstream provider and a hardware failure in the DC-2 switch rack. This combination is rare, but when it happens, the impact is brutal. Think about it: overseas clients can’t reach their servers because the international link is saturated, while local servers in DC-2 are also dropping off because the switch is failing. Hybrid clients — those with servers in DC-2 plus users accessing from overseas — got hit from both sides. Support tickets flooded in — 50+ in the first 30 minutes. But this was exactly the moment our DRC was built for. All those procedures, runbooks, drills we’d done before? Time to execute for real. This wasn’t a drill. This was the real thing. And the team had to move fast but stay calm — panic doesn’t help anyone. Stay focused, prioritize, execute.
DRC Protocol Activation: First Critical Step
The first thing we did was formally activate the DRC protocol. This isn’t just bureaucratic paperwork — it defines who does what, which communication channels to use, the escalation path to management, and the timing for progress reviews. We only had 3 NOC engineers on the night shift — we immediately called in 2 senior team members. Within 10 minutes, we had 5 people on the bridge call, split into two task forces: Team Alpha focused on DDoS mitigation, Team Bravo focused on the switch rack issue. Each had a designated leader and their own documentation. Main communication tools: dedicated Slack channel plus voice bridge to avoid written miscommunication. Status updates every 15 minutes — what’s been done, what’s pending, what’s blocking.
BGP Traffic Rerouting: Changing Lanes Mid-Storm
We have a backup upstream provider — smaller throughput (10 Gbps vs 40 Gbps on the main link) but enough for priority traffic. The problem: by default, BGP routing always prefers the main path because of higher bandwidth. We had to intervene manually. Team Alpha logged into the edge router — a Juniper MX240 — and applied a route-map to set higher local-preference for routes coming from the backup provider. This made outbound traffic use the backup path. But inbound traffic is trickier — we needed to coordinate with the main upstream to advertise our prefixes via the backup provider with a shorter AS path. At the same time, we asked the main upstream to enable RTBH (Remotely Triggered Black Hole) for the identified attack source IPs. The effect? Within 5-10 minutes, inbound traffic dropped from 40 Gbps to around 15 Gbps. Not normal yet, but significantly better. Packet loss dropped from 35% to 8%. Enough for priority clients to start accessing their services again — albeit slowly.
set policy-options policy-statement SET-BACKUP-PREF term 10 from protocol bgp
set policy-options policy-statement SET-BACKUP-PREF term 10 from neighbor 203.0.113.2
set policy-options policy-statement SET-BACKUP-PREF term 10 then local-preference 200
set policy-options policy-statement SET-BACKUP-PREF term 10 then accept
set policy-options policy-statement SET-BACKUP-PREF term 20 then reject
set routing-options router-id 192.0.2.1
set routing-options autonomous-system 65001
Edge Mitigation: FortiGate DDoS Protection Profile
While waiting for the upstream filtering to kick in, we enabled the DDoS protection profile on our FortiGate 600D firewall. We cranked up the settings aggressively — in an emergency situation, better safe than sorry. UDP flood threshold dropped from 5000 pps to 1000 pps per source IP. SYN flood from 3000 to 500. ICMP flood from 2000 to 200. Session limit per source set to 1000 — if any IP creates more than 1000 simultaneous connections, it gets blocked immediately. The side effect? Some legitimate clients with heavy connections could get false positives. But we had a whitelist ready for pre-registered priority client IPs. The comms team sent out notifications to major clients — explaining the situation and asking them to stay patient while we worked on mitigation. Transparency matters. Clients stay calmer when they know what’s happening.
config firewall ddos-policy
edit 1
set service "UDP"
set srcaddr "all"
set dstaddr "all"
set threshold 1000
set action block
next
edit 2
set service "SYN"
set srcaddr "all"
set dstaddr "all"
set threshold 500
set action block
next
edit 3
set service "ICMP"
set srcaddr "all"
set dstaddr "all"
set threshold 200
set action block
next
end
Rack Switch Emergency: The Physical Side of the Problem
While Team Alpha was deep in BGP and firewall configurations, Team Bravo headed to DC-2 for a physical inspection. Results: two fan trays on the main switch in rack 12-C were completely dead. Temperature hit 58°C — dangerously close to the critical threshold of 65°C. If the switch overheated completely, it could shut down and require 30-60 minutes of cooling before restarting. Decision: bypass routing through the backup switch in the same rack, power down the main switch, let it cool for 15 minutes. The backup switch has lower capacity (48 gigabit ports vs 96 10GbE ports), but enough to keep critical servers online. The cutover process took 20 minutes — including cable verification and making sure Spanning Tree was stable. Result: all servers in rack 12-C came back online. Main switch temperature dropped to 42°C after we directed external fans at it. The main switch issue was handled the next business day — fan tray replacement and full diagnostic.
Gradual Recovery: Don’t Rush Traffic Restoration
After about 2 hours of intensive mitigation, the situation was under control. DDoS traffic from the main upstream was being filtered — dropped from 40 Gbps to a residual 2-3 Gbps. The backup provider was handling priority traffic steadily. The switch rack was stable. But we didn’t immediately route everything back to the normal path. The principle: gradual recovery with tight monitoring. Phase 1 (30 mins): restore traffic for critical clients — real-time businesses like payment gateways and trading platforms. Phase 2 (next 30 mins): restore traffic for mid-tier clients — regular VPS and hosting customers. Phase 3 (1 hour): restore all remaining traffic including non-priority. Each phase was reviewed before proceeding — checking packet loss, latency, CPU utilization on the edge router. Any anomaly triggered an immediate rollback. Thankfully, all phases completed smoothly. Total incident response time: 4 hours 20 minutes from first detection to full recovery.
| Metric | Normal | During Attack | After DRC |
|---|---|---|---|
| International Link Utilization | 15-20% | 98-100% | 25% |
| Packet Loss | <0.1% | 35-55% | 0.5% |
| BGP Prefix Count | 250K | 150K (flapping) | 248K |
| Switch Rack Temperature | 37°C | 58°C | 39°C |
| RTT to Singapore | 35ms | 450-500ms | 38ms |
| Active Support Tickets | 0-2 | 50+ | 5 (monitoring) |

Key Lessons from This Incident
After everything settled, we did a full review. A few key takeaways: first, your DRC must be tested regularly — not just a document sitting in a folder. A runbook that’s never been simulated will cause chaos during a real incident. Second, a backup provider isn’t just a nice-to-have — it’s a lifesaver when your main provider goes down. Make sure its capacity is enough for priority traffic, at least 25% of your main capacity. Third, dedicated communication channels for incident response are crucial. Don’t mix incident chat with daily chatter — important info will get buried under memes and stickers. Fourth, proper monitoring alerting with anomaly detection that can distinguish normal traffic from attacks. Alert thresholds shouldn’t be too sensitive (noise) or too loose (you only find out about the fire when the building’s already burning). We immediately wrote a post-mortem and updated several runbooks based on what we learned.
If you manage a datacenter — or even just a handful of servers — please don’t underestimate the importance of a DRC. Incidents like this can happen anytime, to anyone. No matter how good your infrastructure is, without a solid disaster recovery plan that’s been tested, one attack can bring everything crashing down. Investing in DRC isn’t a cost — it’s insurance for your business. Take the time to set up proper server monitoring, make sure your BGP redundancy is working correctly, and keep your firewall configuration up to date. There’s nothing more valuable than hands-on experience — but it’s even better when that experience comes from someone else’s incident, not yours. Learn from others’ mistakes.
Q: How long does recovery take for a massive DDoS incident like this?
Our total incident response time was about 4 hours 20 minutes from first detection to full recovery. But without a well-trained DRC team and tested procedures, this could easily stretch to 12-24 hours — even longer if the switch rack had completely failed. The key factors: preparedness and execution speed.
Q: What’s the difference between regular backups and a DRC?
Regular backups typically store data — if a server fails, you restore from backup. A DRC (Disaster Recovery Center) is much more comprehensive: it covers procedures, alternative infrastructure, a dedicated team, and runbooks for various disaster scenarios including DDoS attacks, hardware failures, natural disasters, and more. DRC is a complete survival plan, not just a data backup strategy.
Q: Was the switch rack failure directly caused by the DDoS attack?
Not directly. The dead fan trays were a pre-existing hardware issue — likely gradual component degradation. But the heavy traffic load passing through that rack during the DDoS accelerated the failure. Think of it like someone with a mild illness being forced to run a marathon — they’ll collapse. The DDoS attack was the indirect trigger that exposed the latent hardware problem.
Q: How do you detect a DDoS attack early?
We use Grafana with netflow data from our edge routers. Our alert threshold is set at 70% of link capacity. But more important than static thresholds is anomaly detection — tools like FastNetMon or Akvorado that learn normal traffic patterns and detect abnormal spikes in real-time. The combination of both approaches gives you: static thresholds for baseline monitoring + anomaly detection for early warnings when something unusual starts happening.
So here’s the deal — incidents like this are unavoidable in our line of work. But solid preparation can be the difference between a full-blown disaster and a manageable incident. Take the lessons from this story, audit your own DRC setup, and make sure your team is ready to execute when things go south. Ever dealt with a nastier DDoS attack? Got a different mitigation approach that worked better for you? Drop it in the comments — you never know who might benefit from your experience. Thanks for reading, stay safe out there, and may your links never saturate!