Sticky Sessions with Apache APISIX: Session Affinity Guide

Nicolas Fränkel

Nicolas Fränkel

July 27, 2023

Technology

Sticky sessions, also known as session affinity, route requests carrying the same session identifier to the same upstream node. They can help a stateful application reuse node-local session data, but they do not make that data durable. This guide explains the trade-offs and shows how to configure cookie-based affinity with Apache APISIX.

Why Sticky Sessions?

Sticky sessions became popular when we stored the state on the upstream node, not the database. I'll use the example of a simplified e-commerce shop to explain further.

The basic foundations of a small e-commerce site can consist of a web application and a database.

Foundation architecture of an e-commerce app

If the business is successful, it will grow, and you'll need to scale this architecture at some point. Once you cannot scale vertically (bigger machines), you must scale horizontally (more nodes). With additional app nodes, you'll also need a load balancer mechanism in front of the web app nodes to distribute the load among them.

Load Balancing architecture for additional nodes

Going to the database every time is an expensive operation. It's okay for data that is accessed infrequently. However, we want to display the cart's content for every request. A couple of alternatives are available to speed things up. If we assume that the web app uses Server-Side Rendering, the classical solution is to keep cart-related data in memory on the web app node.

However, if we store user X's cart on node 1, we need to ensure that we forward every request of user X to the same node. Otherwise, they will feel as if they lost their cart's content. Sticky sessions, or session affinity, is the mechanism that consistently routes the same user to the same node.

Limitation of Sticky Sessions

Before going further, consider a significant limitation: if a web application node is the only place that stores session data and that node fails, the data are lost. In the e-commerce example, users could lose their carts.

Do not treat affinity as a substitute for state durability. Depending on the application, use session replication, a shared session store, or a stateless design. These approaches let another node continue serving the session when the preferred node is unavailable.

While session replication exists in all tech stacks, there's no related specification. I'm familiar with the JVM, so here are a couple of options:

When data is replicated on all nodes (or a remote cluster), you may think you no longer need sticky sessions. It's true if one accounts only for availability and not for performance. It's about data locality: fetching data on the current node rather than from somewhere else via the network is faster.

Sticky Sessions on Apache APISIX

For applications that benefit from session affinity, Apache APISIX can combine a consistent-hash upstream with a stable request value, such as an application session cookie. Use this approach only when the application accepts the durability and load-distribution trade-offs described above.

Apache APISIX binds a route to an upstream. An upstream consists of one or more nodes. When a request matches the route, Apache APISIX must choose among all available nodes to forward the request to. By default, the algorithm is weighted round-robin. Round-robin uses one node after the other, and after the last one, get back to the first one. With a weighted round-robin, the weight affects how many requests Apache APISIX forwards to a node before it switches to the next one.

However, other algorithms are available:

  • Consistent hashing
  • Exponentially weighted moving average (EWMA)
  • Least connection
  • A custom-made one

Consistent hashing allows forwarding to the same node depending on some value: an NGINX variable, an HTTP header, a cookie, etc.

HTTP is stateless, so many application servers set a cookie to associate later requests with a session. To hash on that cookie, you need its exact name. Common examples include:

  • JSESSIONID for JVM-based servers
  • PHPSESSID for PHP
  • ASPSESSIONID for ASP.Net
  • etc.

I shall use a regular Tomcat, so the session cookie is JSESSIONID. Henceforth, the Apache APISIX documentation for two nodes is the following:

routes: - uri: /* upstream: nodes: "tomcat1:8080": 1 #1 "tomcat2:8080": 1 #1 type: chash #2 hash_on: cookie #3 key: JSESSIONID #4
  1. Define the upstream nodes
  2. Choose the consistent hashing algorithm
  3. Hash on cookie
  4. Define the case-sensitive cookie name to hash on

APISIX expects the cookie name in key when hash_on is cookie. In APISIX 3.18, if a request does not contain that cookie, APISIX falls back to remote_addr as the hash key. The initial request can therefore reach one node, while the next request can reach another after the application sets JSESSIONID. Applications that need continuity from the first request should establish a stable cookie before creating node-local session state or use shared or replicated session storage.

The configuration is equivalent to using the NGINX variable for that cookie internally. Once requests carry the cookie, expose a test response that identifies the serving node, or inspect the application logs, then send repeated requests with the same cookie value:

for request in 1 2 3; do curl -s http://127.0.0.1:9080/ \ -H 'Cookie: JSESSIONID=test-session' echo done

All three requests should identify the same healthy upstream. Then verify that both upstreams are healthy and sample several distinct cookie values:

for session in $(seq 1 20); do printf 'session-%s: ' "$session" curl -s http://127.0.0.1:9080/ \ -H "Cookie: JSESSIONID=session-$session" echo done

Different cookie values can still hash to the same node. The goal is to confirm stable routing for each repeated key and, across a sufficiently large sample, observe that more than one healthy upstream participates.

When upstream health checks mark the selected node unhealthy, APISIX can choose from the remaining healthy nodes. That failover does not recover session data stored only on the failed node.

Conclusion

Sticky sessions can reduce remote session lookups and support applications that still keep state on individual nodes. Use them only with an explicit durability and failover strategy, then monitor load distribution because affinity can concentrate traffic unevenly.

To go further:

Tags:
Share article link