Application and service failures

Listen to this lesson

Episode 43 · 45:48

Every episode of this course is also a podcast: listen on Spotify.

This episode is a study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.

Objective 4.2 · Troubleshooting · 28% of the exam

Why this matters

Most of the time, the hardware is fine and the operating system is running, and still something users depend on is not working. The web application returns errors, the database will not accept connections, the backup agent has stopped. From the users' point of view the server is down; from the server's point of view everything looks normal.

Application and service faults have a small set of common causes, and most can be found quickly with the right questions. This lesson covers services that will not start, the accounts and permissions they run under, two programs fighting over the same port, servers running out of memory or disk, and the application's own logs, which usually name the problem outright.

The lesson

Services that will not start, and their dependencies

Server applications usually run as services on Windows or daemons managed by systemd on Linux, starting automatically at boot and running in the background.

When a service will not start:

  • Check its status and the error. On Windows, the Services console and Get-Service show the state, and the System log records service failures, often with an error code. On Linux, systemctl status followed by the service name shows the state, the last few log lines and the exit code, and journalctl -u shows its full log.
  • Check its dependencies. Services often depend on other services: a web application on its database, a directory-integrated service on the network and the domain. A dependency that has failed or not yet started stops the service that needs it. On Windows, the service's properties list its dependencies; on Linux, the unit file declares them with Requires and After.
  • Check its configuration. A service that stopped starting after a change to its configuration file very often has a typing error in it. Many applications have a way to check configuration without starting, such as nginx -t or apachectl configtest.
  • Check the startup type. A service set to Manual or Disabled will not start at boot, and a service disabled on purpose may have been disabled for a reason.

Configure recovery options so that services restart automatically after a failure, but monitor restarts: a service restarting every few minutes has a problem that automatic restart is hiding.

Permissions and service account problems

Every service runs as some account, and a large share of service failures come down to that account.

  • Expired or changed passwords. A service configured with a domain account and its password will fail to start, typically with a logon failure error, when that password expires or is changed without updating the service. This is one reason the accounts lesson recommended managed service accounts, whose passwords the domain rotates automatically.
  • Locked or disabled accounts. A service account locked out by failed logons, perhaps from an old copy of the password somewhere else, or disabled by an access review that did not know it was in use.
  • Missing rights. The account may lack the right to log on as a service, or permission to read its program files, write to its log or data folders, or reach a network share or database.
  • Permissions changed underneath it. A hardening change or a tidy-up of folder permissions can remove access a service relied on.

On Linux, check the owner and permissions of the files the service needs with ls -l, and remember that mandatory access control systems such as SELinux can block access that ordinary permissions allow. SELinux denials appear in the audit log, and a file copied into place with the wrong security context is a classic cause.

The error message usually says "access denied" or "permission denied" somewhere. Find out which account the service runs as, and which resource it could not reach.

Port conflicts

Two programs cannot listen on the same port at the same address. If a second one tries, it fails to start, usually with an error that the address is already in use.

Port conflicts typically arise when:

  • a second web server or application is installed that also wants port 80 or 443;
  • a service is started twice, or an old copy has not fully stopped;
  • an application's port is changed to one already in use by something else;
  • a development tool or agent takes a port the production service needs.

To find what is using a port:

  • On Windows, netstat -ano lists listening ports with the process ID of each, which Task Manager or Get-Process then names; Get-NetTCPConnection does the same in PowerShell.
  • On Linux, ss -tulpn lists listening sockets with the process name.

Resolve it by stopping or reconfiguring the program that should not be there, or moving one service to a different port and updating anything that connects to it, including firewall rules and load balancer settings. Remember the host firewall too: a service that starts and listens correctly can still be unreachable if the firewall blocks its port, which the attack surface lesson set to default deny.

Resource exhaustion: memory leaks and full disks

A service may fail, not because anything is wrong with it, but because the server has run out of something it needs.

Memory leaks. A program with a memory leak fails to release memory it no longer needs, so its usage grows steadily until the server runs short. The symptom is a server that works after a restart and degrades over days or weeks: slower responses, heavy paging, then failures. Monitoring shows one process's memory climbing without levelling off. On Linux, the kernel's out-of-memory killer may end processes to free memory, recorded in the kernel log. Restarting the service buys time; the real fix is an update from the vendor or developer.

Full disks. When a volume fills up, applications cannot write their data, logs or temporary files, and they fail, sometimes in ways that do not obviously mention disk space. Databases may stop accepting writes; services may fail to start; logs may simply stop. Common culprits are logs that are never rotated, old backups and dump files, temporary files, and on Linux, a small /var or root partition. Check free space with df -h on Linux or in File Explorer and Get-Volume on Windows, and find what filled it with du on Linux or a disk usage tool on Windows.

A related Linux trap is running out of inodes, the file system's records for files: a disk with free space but millions of tiny files can refuse to create new ones. df -i shows inode usage.

Both problems are best caught by the monitoring and alert thresholds from the monitoring lesson, before they cause a failure.

Reading application logs

Many applications write their own logs, separate from the system log, and these are usually the most specific source of information about what went wrong.

  • On Windows, some applications write to the Application log in Event Viewer, and many keep their own log files, for example under the application's installation folder or ProgramData. IIS keeps its logs in the logs folder under inetpub.
  • On Linux, application logs are usually in /var/log, often in a directory named for the application, such as /var/log/nginx, or in the journal for services run by systemd.
  • Databases, web servers and application platforms each have their own error logs, separate from access logs.

When reading them:

  • find the time the problem started and read around it;
  • look for the first error, since one failure often causes a cascade of others that are only symptoms;
  • search the exact error message in the vendor's documentation and knowledge base;
  • raise the log level temporarily if the logs do not say enough, and remember to lower it again, since verbose logging can itself fill a disk.

Correlate application logs with the system log and monitoring data. A database error at 02:14 means more when the system log shows the disk filled at 02:13.

Try it

An interactive exercise runs here: a real Linux machine in your browser that checks each step. The commands above work on any Linux machine too.

Practise what you just read

1. A web server fails to start with 'Address already in use'. What should the administrator do first?

Select one

  1. Change the server's IP address to free up the port
  2. Reinstall the web server, as its files are corrupted
  3. Find which process holds the port with netstat or ss
  4. Restart the network switch the server is connected to
Show answer

C. Another program is listening on the web server's port. netstat -ano on Windows or ss -tlnp on Linux shows which process, which can then be stopped or reconfigured.

2. A service configured with a domain account fails to start with a logon failure after the account's password was changed. What is wrong?

Select one

  1. The service still holds the old password
  2. The password change deleted the service account
  3. The server needs a new service licence
  4. The firewall is blocking the service's logon
Show answer

A. Services configured with an account and password must be updated when the password changes. Managed service accounts avoid this because the domain rotates the password itself.

3. A server works well after each restart but grows slower over two weeks, with one process's memory rising steadily. What is the likely cause?

Select one

  1. A DNS fault
  2. A duplex mismatch
  3. Disk fragmentation
  4. A memory leak
Show answer

D. Memory that grows without levelling off is the signature of a leak. Restarting buys time; the real fix is an update from the vendor or developer.

7 more questions on this objective are part of the full course.

Practise the full question bank in the exam simulator

Hands-on labs

All hands-on labs

This is an independent study companion for CompTIA Server+ SK0-005 and is not produced by or endorsed by CompTIA.