> ## Content Index
> Fetch the complete content index at: https://www.bitsinflight.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Networking Needs More Telemetry
- URL: https://www.bitsinflight.com/networking-needs-more-telemetry-tfd/
- Published: 2026-04-07T12:00:00.000Z
- Updated: 2026-09-16T16:01:10.000Z
- Description: A Tech Field Day Podcast panel ahead of NFD40 with Tom Hollingsworth and Scott Robohn. Modern gear can finally give us all the telemetry we want. The harder problem is what to do with it.
- Author: Jason Gintert
- Tags: media, podcast, video, NFD, nfd40, AI

*Tech Field Day Podcast, recorded ahead of Networking Field Day 40\. Published 7 April 2026\. Host: Tom Hollingsworth, with fellow delegates Scott Robohn and Pete.*

The premise Tom put to us was that networking needs more telemetry. Nobody argued with that. The interesting part was everything that follows from it.

## Why now

Scott's framing: a long time ago nobody cared about real-time response from the network, and 30-second RIP updates were fine. Now the IP network carries everything that matters, much of it sensitive to human expectations, so we have to measure and respond far faster than we did twenty years ago. Pete added that network engineers have become the jacks-of-all-trades. We are the people who understand flows across storage, Wi-Fi and applications, and without broad visibility we just get strange tickets we cannot explain.

My angle was that the gear can finally do it. Early in my career I could drop a router with overzealous SNMP polling, and I have the bruises to prove it. Modern equipment has dedicated ASICs and the raw horsepower to produce telemetry without falling over. The hyperscalers needed that to run their enormous architectures, and those capabilities are now trickling down to the rest of us.

## There is such a thing as too much data

Everyone eventually says the word AI, and we got there about twenty minutes in. The promise is sifting through gobs of telemetry. The caveat is context: you can overload an LLM just as easily as a human, so all of this has to be distilled and structured before it is useful to anyone. Scott made the important distinction that none of this lives on the routers. Systems like Kentik and Selector pull it into one place so LLMs, machine learning and plain old statistical regression can work on it together. His phrase for the job was finding the needle in the needle stack.

Tom raised the old sampling problem, now with more zeros: with AI clusters where tail latency costs real money, you cannot act on data that is four seconds old. Scott's answer was that real time has to happen on chip and the rest of us are back-seat driving, looking for the smoking gun that says we need more capacity or a different design.

## Not every problem is an F1 pit stop

Tom's best point was about immediacy bias. People want the dashboard to show them the problem right now so they can fix it right now, and that is how you end up upgrading the wrong link and chasing the rabbit through the whole network. The value of more telemetry is that it lets you find the actual cause before you fix the wrong symptom. I mentioned the overnight use case I keep seeing, where an LLM crunches the night's data and the operator walks in to a summary rather than a mess.

## One piece of advice each

Asked for a single recommendation, mine was to look at the system holistically. Application performance, security and network performance have been treated as separate domains for a long time, and AI is bringing tools that can finally read them together. Pete's was to make contacts outside the networking group and learn enough of their language to cross the boundary. And I added the unglamorous one: make sure the data is clean. I have been in environments where flows were double-counted and every number was wrong. Trust, but verify.

[Watch on YouTube →](https://www.youtube.com/watch?v=XWSau3rpHFA&ref=bitsinflight.com)