Skip to content
Let’s plan the right software for your processes. Call us for a demo or a quote: +90 546 737 48 29

TR EN DE

What is load balancing?

Load balancing means sharing the requests that reach a website or application among several servers that do the same job. The visitor connects to a single address; the load balancer that receives the request picks one of the servers behind it and passes the request on. The visitor does not know how many servers there are, and does not need to.

Think of the checkouts in a supermarket: with one checkout open the queue grows, and if that checkout breaks down, sales stop. With several checkouts and a member of staff directing customers to a free one, the queue is shorter and one checkout closing does not bring everything to a halt. This guide first explains what load balancing gives you, then how it is defined in Nginx and what you need to watch out for.

In brief

  • Load balancing spreads incoming requests across several servers running the same application.
  • In Nginx the server group is defined with an upstream block; the default method is round-robin.
  • Health checks in open-source Nginx are passive: they work through max_fails and fail_timeout.
  • Sessions, uploaded files and the database have to be planned from the start.

Note

The IP addresses and domain names in the examples are placeholders. After changing the configuration, always test it first and then reload; file paths and service names vary by distribution.

On this page

How does load balancing work?

  1. Requests are spread across several servers

    A load balancer is a reverse proxy that sits between the visitor and the backend servers. It brings three main benefits:

    • Capacity: A load that one server cannot handle is divided among several. When demand grows, a new server is added to the group.
    • Availability: If one of the servers fails, requests are sent to the remaining ones; the site may slow down but it stays up.
    • Easier maintenance: The server to be updated is taken out of the group and put back when the work is done. Visitors do not notice the maintenance.

    Load balancing is not a speed setting in itself: it does not make a slow query faster, it only allows more requests to be served at the same time.

    Diagram: visitors' requests reach the load balancer and are distributed to three backend servers : Enlarge
  2. Choose a balancing method

    The balancing method decides which server gets the next request. The three methods used most often in Nginx are:

    MethodHow does it distribute?When is it suitable?
    round-robin (default)Hands out requests in turn, according to the weights.When the servers are of similar power and requests take a similar time.
    least_connGives the request to the server with the fewest active connections at that moment.When request durations differ widely (long reports, file downloads).
    ip_hashSends requests from the same client address to the same server.As a stopgap when sessions are kept on the server.

    If in doubt, start with the default method; it is enough for most sites.

    Diagram: comparison of the round-robin, least_conn and ip_hash balancing methods : Enlarge
  3. A server that stops responding is taken out of rotation

    A load balancer must be able to stop sending requests to a faulty server. This is called a health check. In open-source Nginx the check is passive: Nginx does not probe the servers separately, it looks at the outcome of real visitor requests. If communication with a server fails a certain number of times within a certain period, that server is not used for a while; once the period is over it is tried again.

    Active health checks, which probe the servers at regular intervals of their own accord, are not built into the open-source version; they are found in the commercial version, in add-on modules or in other load balancer software.

    Flow diagram: a request fails, the failure count reaches the limit, the server is not used for a while, and it is tried again once the time is up : Enlarge
  4. Plan for sessions

    Many applications keep session data (login state, basket) on the server's own disk. With one server that is not a problem; with several, the visitor's first request may go to server A and the second to server B, which does not recognise them. The result: the user is suddenly logged out or the basket is empty. This is known as the session stickiness problem; the solutions are covered in a separate section below.

    Diagram: three solutions to the session problem; ip_hash, a shared session store and a stateless identity token : Enlarge

The upstream block in Nginx

In Nginx the backend servers are defined as a group with an upstream block. The block goes inside the http block, outside any server block. The group is given a name, and that name is used in the proxy_pass directive:

Nginx
# Inside the http block: the group of backend servers
upstream app_backend {
    server 10.0.0.11:8080;
    server 10.0.0.12:8080;
    server 10.0.0.13:8080;
}

server {
    listen 80;
    server_name example.com www.example.com;

    location / {
        proxy_pass http://app_backend;
        proxy_set_header Host $host;
        proxy_set_header X-Real-IP $remote_addr;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        proxy_set_header X-Forwarded-Proto $scheme;
    }
}

As no other setting is made in this example, requests are handed to the three servers in turn. The proxy_set_header lines let the backend server know the domain name the visitor asked for, the real IP address and whether the connection was HTTPS; without them the application believes every request comes from the load balancer. For the basic set-up, see the reverse proxy set-up guide.

If you use HTTPS, the certificate is normally installed on the load balancer (SSL/TLS termination); the details are in the HTTPS in Nginx guide.

After every change, test the configuration first and reload only if there are no errors:

Bash
# Test the configuration
sudo nginx -t
# If there are no errors, reload without dropping connections
sudo systemctl reload nginx

Methods: round-robin, least_conn, ip_hash

round-robin needs no directive; it is the default. The other methods are added as a single line at the top of the upstream block.

least_conn gives the request to the server with the fewest active connections at that moment; the weights are taken into account as well:

Nginx
upstream app_backend {
    least_conn;
    server 10.0.0.11:8080;
    server 10.0.0.12:8080;
    server 10.0.0.13:8080;
}

ip_hash picks the server from the client's IP address; requests from the same address always go to the same server for as long as that server is available. For IPv4 addresses the first three octets are used, for IPv6 the whole address.

Nginx
upstream app_backend {
    ip_hash;
    server 10.0.0.11:8080;
    server 10.0.0.12:8080 down;   # under maintenance: mappings are preserved
    server 10.0.0.13:8080;
}

If you need to remove a server temporarily while using ip_hash, mark it with down instead of deleting the line; that way the server mapping of the other clients is preserved.

The weight, backup and down parameters

Parameters that control a server's behaviour can be added at the end of each server line:

  • weight: The server's weight; the default is 1. A server with weight 3 receives roughly three times as many requests as a server with weight 1. Use it to give a more powerful server a larger share.
  • backup: A backup server. It receives requests only when none of the primary servers is available. It cannot be combined with the ip_hash method.
  • down: Marks the server as permanently out of use. It is the simplest way to take a server out of the group during maintenance.
Nginx
upstream app_backend {
    server 10.0.0.11:8080 weight=3;   # more powerful server: more requests
    server 10.0.0.12:8080;            # weight=1 (default)
    server 10.0.0.13:8080 down;       # under maintenance, receives no requests
    server 10.0.0.14:8080 backup;     # only if the others are unavailable
}

The maintenance routine is: add down to the server's line, test and reload the configuration, carry out the maintenance, remove the word down and reload again. Connections in progress are not cut during a reload; the old worker processes finish the requests they are handling and then exit.

Health checks: max_fails and fail_timeout

Passive health checks are tuned with two parameters:

  • max_fails: The number of failed attempts needed for the server to be considered unavailable. The default is 1; setting it to 0 switches the counting off.
  • fail_timeout: Both the period in which failed attempts are counted and the time for which the server is considered unavailable. The default is 10 seconds.
Nginx
upstream app_backend {
    # 3 failed attempts in 30 seconds: server is not used for 30 seconds
    server 10.0.0.11:8080 max_fails=3 fail_timeout=30s;
    server 10.0.0.12:8080 max_fails=3 fail_timeout=30s;
    server 10.0.0.13:8080 max_fails=3 fail_timeout=30s;
}

server {
    listen 80;
    server_name example.com;

    location / {
        proxy_pass http://app_backend;
        proxy_set_header Host $host;
        proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
        # Do not wait long for a server that is down
        proxy_connect_timeout 3s;
        # In which cases should the next server be tried?
        proxy_next_upstream error timeout http_502 http_503;
        proxy_next_upstream_tries 2;
    }
}

In this example, if communication with a server fails 3 times within 30 seconds, no requests are sent to that server for 30 seconds. When the time is up, Nginx tries the server again with a real request.

What counts as a "failure" is set by the proxy_next_upstream directive. By default a connection error and a timeout count as failures; in the example, 502 and 503 responses have been added. The failed request is passed to the next server; proxy_next_upstream_tries limits the number of such attempts. Keeping proxy_connect_timeout short prevents a server that is down from keeping the visitor waiting.

The limit of passive checking is this: a fault is noticed only when a visitor's request fails. So monitor the servers from outside as well; the uptime monitoring advice in the backup and monitoring guide applies here too. For the causes of the 502 and 504 errors visitors may see, read the 502 and 504 errors guide.

Session stickiness and how to solve it

If session data sits on the disk or in the memory of a single server, every request from that visitor has to reach that server. There are three ways to deal with this:

  • Routing to the same server with ip_hash: The quickest fix to apply, but a limited one. Many users leaving the same network (a company or an institution) pile up on one server; a mobile user loses the session when their IP address changes; when a server drops out, the sessions on it are lost. If a CDN or another proxy sits in front of the load balancer, Nginx sees the address of that layer and not the visitor's, and the distribution becomes lopsided.
  • Keeping sessions in a shared store: Sessions are not kept on the servers' disks but in a common place that all of them can reach (an in-memory store such as Redis or Memcached, or the database). Whichever server answers, the session is found. This is usually the lasting solution.
  • A stateless identity token: The server stores no session; the user's identity travels in a token that the server has signed and that is sent with every request. As every server can verify the signature, the request is accepted wherever it lands. The signing key must be the same on all servers and the token's lifetime should be kept short.

For secure settings for session cookies, see the login and session security guide.

Things to watch out for

  • Uploaded files must be shared: If an image uploaded by a user is written only to the disk of the server that took the request, the other servers cannot find it. Keep uploaded files in shared storage that all servers can reach, or synchronise them between the servers.
  • The database remains a single point: Multiplying the application servers does not multiply the database. All servers connect to the same database; if the slowness or the fault is there, load balancing will not fix it. The database needs its own backup and standby plan.
  • The load balancer itself is a single point of failure: Even with three servers behind it, if the one load balancer in front stops, the site goes down. Where uninterrupted service matters, the load balancer is made redundant too (a second load balancer and a shared IP address that moves to it on failure), or the hosting provider's managed load balancing service is used.
  • Releases must be deployed consistently: If the servers run different versions of the code, the same visitor sees the new page on one request and the old one on the next. Deploy by taking the servers out of the group one at a time, and make sure database changes also work with the old version.
  • Scheduled tasks: A scheduled task that runs separately on every server does the same job several times (for example, sends the same email three times). Run such tasks on one server only, or use a lock.
  • Cache and logs: A cache kept per server can give different results from one server to the next. Logs are scattered as well; when looking for a problem you have to check all of them.

If you are not sure that one server is no longer enough, consider improving the existing server first: caching and compression often bring more gain for less effort.

Frequently asked questions

How many servers do I need for load balancing?

You need at least two application servers behind it; the load balancer runs on a separate server or as a managed service. With a single application server, load balancing brings no benefit.

Does load balancing make my site faster?

It does not shorten the time a single request takes. It allows more requests to be served at once; you will notice it on a site that slows down under heavy traffic, but it does not fix a slow query or a heavy page.

Does open-source Nginx notice a faulty server by itself?

Yes, but passively: it looks at real requests failing (max_fails and fail_timeout). Active health checks that probe the servers at regular intervals are not built into the open-source version.

Does ip_hash solve the session problem for good?

No. It is a quick stopgap; the session is lost when the server drops out or the visitor's IP address changes. The lasting solution is to keep sessions in a shared store or to use a stateless identity token.

Is a load balancer the same thing as a CDN?

No. A load balancer distributes requests among your own servers; a CDN serves content from points close to the visitor and filters traffic before it reaches your server. The two can be used together.

BYK Yazılım Support Team
This guide is written and regularly reviewed by the BYK Yazılım support team. Last updated: 4 October 2026.

Related guides

Let us talk about your website infrastructure

BYK Yazılım builds corporate websites. Write to us with any questions about your site.

Contact us Our corporate website service