Network latency and slow services

Listen to this lesson

Episode 45 · 48:27

Every episode of this course is also a podcast: listen on Spotify.

This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.

Objective 4.3 · Troubleshooting · 28% of the exam

Why this matters

"It's slow" is the most common and least precise complaint an administrator hears. A service can be slow because the server is overloaded, because the application is waiting on a database, because the network between user and server is congested, or because every connection starts with a name lookup that takes five seconds. Each of those has a different owner and a different fix, and the first job is working out which one it is.

This lesson covers the three network measurements that are easily confused, the tools that measure them along a path, congestion and the priority systems that manage it, the name resolution delay that looks like a slow service, and how to decide whether the server or the network is to blame.

The lesson

Latency, bandwidth and packet loss, and which one is the problem

Three different properties of a network connection are all described as "slow", and they have different causes and fixes.

  • Latency is the time a packet takes to travel from one end to the other, usually measured as a round-trip time in milliseconds. It depends on distance, the number of devices on the path, and how long packets wait in queues. High latency hurts interactive and chatty applications, those that exchange many small messages back and forth, since every exchange waits a full round trip.
  • Bandwidth is how much data a link can carry per second; throughput is how much it actually carries. Low bandwidth hurts large transfers, such as backups, file copies and video, while small transactions barely notice it.
  • Packet loss is the percentage of packets that never arrive. Even a small amount, one or two per cent, badly hurts TCP performance, because TCP retransmits lost data and slows down in response to loss. Loss usually comes from congestion, faulty cabling or hardware, or wireless interference.

Jitter, variation in latency from packet to packet, matters mostly for real-time traffic such as voice and video.

The distinction guides the fix. More bandwidth does nothing for a problem caused by latency; a file transfer between two continents can be slow on a fast link simply because of the round-trip time. Identify which property is poor before deciding what to change.

Measuring with ping, traceroute, pathping and mtr

Four tools measure a path through the network.

ping sends ICMP echo requests and reports whether replies came back and the round-trip time of each. It shows latency, and, over many packets, loss. Run it for a while rather than four packets: ping -t on Windows continues until stopped, and ping -c 100 on Linux sends a hundred. Compare against a baseline; 40 ms may be normal to a distant site and alarming to a server in the same room.

traceroute (tracert on Windows) lists each hop, each router along the path, with the latency to it. It shows where latency increases, and where a path stops if packets are not getting through.

pathping, on Windows, combines the two: it traces the route, then sends many packets to every hop over a period and reports loss at each hop.

mtr, on Linux, does the same continuously, updating a live table of latency and loss for every hop, which makes it the preferred tool for intermittent problems.

Two cautions when reading them:

  • Many routers give low priority to answering ICMP themselves, or do not answer at all, so a single hop showing loss or high latency while later hops are fine is usually that router ignoring ping, not a real problem. Real loss or latency appears at a hop and carries on to every hop after it.
  • Firewalls often block ICMP, so a failed ping does not prove a host is down. Test the actual service port instead, for example with Test-NetConnection on Windows or nc or curl on Linux.

Congestion and quality of service

Congestion happens when more traffic is sent across a link than it can carry. Devices queue the extra packets, which adds latency, and when queues are full they drop packets, causing loss. Congestion is often periodic, appearing at the same times each day: the start of the working day, a nightly backup running across a link shared with users, or a large file transfer.

Signs of congestion include latency that rises and falls with load, loss at a particular link, and interface utilisation near 100 per cent on a switch or router, visible in monitoring. Interface errors, by contrast, point to physical problems, covered in the connectivity lesson.

The fixes are to reduce traffic, move it, or manage it:

  • Schedule heavy traffic such as backups and replication outside busy periods, or on separate links or networks.
  • Add capacity where a link is genuinely too small.
  • Use quality of service (QoS) to decide which traffic goes first when a link is congested. QoS classifies traffic and marks it, commonly with DSCP values in the IP header, so network devices can place important or time-sensitive traffic, such as voice and critical applications, in priority queues, and limit less important traffic, such as backups.

QoS does not create bandwidth. It decides who waits when there is not enough, so it helps with short periods of congestion but cannot fix a link that is constantly overloaded.

Slow name resolution disguised as a slow service

Some "slow services" are not slow at all. The service answers instantly, but something before the connection takes several seconds, most often name resolution.

The symptom is a fixed delay, often a few seconds, at the start of each connection, with the service fast once connected. Common causes:

  • the first DNS server configured on the client or server is unreachable, so each lookup waits for it to time out before trying the second;
  • a service performing reverse DNS lookups on each connecting client, such as SSH or a web server's logging, where the reverse zone is missing or unreachable, as the DNS lesson described, so every connection waits on a failed lookup;
  • an application trying IPv6 first, to an address that does not work, before falling back to IPv4;
  • a slow or overloaded DNS server.

To test, time a lookup directly with nslookup, dig or Resolve-DnsName, and compare connecting by IP address with connecting by name. If the IP address connects instantly and the name does not, the delay is in name resolution, and the fix belongs in DNS configuration, not the application server.

Separating a slow server from a slow network

The key question is whether time is being spent on the network or on the server.

Evidence that the network is responsible:

  • ping and traceroute show high latency or loss to the server;
  • the problem affects many services at one location, or everyone on one link;
  • users at other locations, or on the same network segment as the server, find it fast;
  • large transfers are slow while the server's resources are idle.

Evidence that the server is responsible:

  • network measurements are normal: low latency and no loss;
  • the server's resources are strained compared with its baseline, from the monitoring lesson: high processor use, memory pressure and paging, or high disk latency;
  • the application is slow even when tested from the server itself, which removes the network from the path entirely;
  • application logs show slow database queries or waits on another back-end service.

Useful comparisons isolate one variable at a time: test from the server's own console, from a machine on the same switch, and from the user's location. Measure where the time goes in the application too, since many are slow because a back-end service they depend on is slow. The time spent narrowing the problem down is repaid by fixing the right thing, and handing the network team a problem that is theirs, with evidence.

Try it

An interactive exercise runs here: a real Linux machine in your browser that checks each step. The commands above work on any Linux machine too.

Practise what you just read

1. A large file transfer between two continents is slow on a fast link, while small web requests are only slightly delayed. What is the main factor?

Select one

  1. Latency from the round-trip distance
  2. Insufficient bandwidth for large files
  3. A DNS misconfiguration at one end
  4. A duplex mismatch on the server
Show answer

A. Long distances add round-trip time, and TCP transfers wait on acknowledgements, so throughput for one connection falls even on a fast link. Adding bandwidth does nothing for latency.

2. A service responds instantly once connected, but every new connection starts with a fixed five-second pause. What is the likely cause?

Select one

  1. The server's system disk is full
  2. Packet loss on the server's link
  3. Slow or failing name resolution
  4. The application needs more memory
Show answer

C. A fixed delay before each connection, with fast service afterwards, is typical of an unreachable first DNS server timing out, or a failed reverse lookup. Connecting by IP address confirms it.

3. An mtr trace shows 40 per cent loss at hop 4, but 0 per cent at hops 5 to 9. What does this mean?

Select one

  1. The source server's network adapter is faulty and dropping packets
  2. Hop 4 is dropping 40 per cent of all traffic passing through it
  3. The destination is unreachable, so the trace stops at hop 4
  4. Hop 4 is deprioritising replies to probes; there is no real loss
Show answer

D. Real loss continues through every later hop. Loss that appears at one router and vanishes after it means that router gives low priority to answering probes itself, not that it drops forwarded traffic.

7 more questions on this objective are part of the full course.

Practise the full question bank in the exam simulator

Hands-on labs

All hands-on labs

This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.