At langgraph-checkpoint-postgres 3.1.2 (2026-08-07), PostgresSaver.from_conn_string and its async twin open one psycopg connection with autocommit=True, prepare_threshold=0 and dict_row, hold it for as long as the context manager is open, usually the life of the process, and accept no pool, keepalive or prepare_threshold argument. Both savers also set supports_pipeline from the client libpq version alone, so on any libpq 14 or newer every checkpoint write enters pipeline mode, and passing pipeline=False to the factory does not change that. Behind PgBouncer or another transaction-mode pooler, prepare_threshold=0 is the most aggressive setting psycopg offers: every statement is server-side prepared on first use. Four upstream issues about this path are open (#3716, #5675, #7304, #8420), and three pull requests that would have added knobs were closed without merge. The fix sits entirely on the client: build an AsyncConnectionPool with autocommit, dict_row, prepare_threshold=None and check=AsyncConnectionPool.check_connection, pass the pool to AsyncPostgresSaver, and set supports_pipeline = False when a pooler is in the path. Retention is your job too: delete_thread is the only deletion API, there is no TTL and no timestamp column. Start by grepping your code for from_conn_string.
Every LangGraph checkpointer Postgres tutorial works on the first try, on a laptop. The trouble starts when the same code runs against the bank's database through PgBouncer and a firewall that reaps idle sessions, and checkpoint writes begin failing with SSL error: bad length, SSL connection has been closed unexpectedly or the connection is closed, intermittently, and not in staging. The official LangGraph page on adding memory shows one construction: PostgresSaver.from_conn_string(DB_URI) with sslmode=disable on localhost. That call is where most of the production behaviour is decided, and you cannot tune any of it through the call itself.
At langgraph-checkpoint-postgres 3.1.2, published 2026-08-07 and still the latest release on the date of this post, from_conn_string opens one psycopg connection with prepare_threshold=0, holds it for as long as the context manager is open, and on any libpq 14 or newer still runs every checkpoint write in pipeline mode when you pass pipeline=False. Four upstream issues about this path are open: #3716, #5675, #7304 and #8420. Three pull requests that would have added the missing knobs, #7311, #8020 and #8421, were closed without merge.
Every claim below was read against langgraph-checkpoint-postgres 3.1.2, psycopg 3.3.6 and psycopg-pool 3.3. Nothing here was run by us; it is source and documentation. If you are still choosing a framework, start with the LangGraph, CrewAI and OpenAI Agents SDK comparison.
What from_conn_string actually does in 3.1.2
The sync factory lives in langgraph/checkpoint/postgres/__init__.py.
Connection.connect(conn_string, autocommit=True, prepare_threshold=0, row_factory=dict_row).pipeline=True it yields cls(conn, pipe) inside conn.pipeline(); otherwise it yields cls(conn).pipeline. The async version in aio.py adds serde and nothing else.There is no pool, no pool_config, no keepalive argument (libpq keepalive parameters can still ride in the connection string), no prepare_threshold and no way to change pipeline behaviour. Issue #7304 asked for pool_config on AsyncPostgresSaver.from_conn_string(); it is open, and the PR that implemented it, #7311, was closed without merge on 2026-03-27.
The constructors are more flexible than the factory. PostgresSaver(conn, pipe=None, serde=None) and AsyncPostgresSaver(conn, pipe=None, serde=None) accept either a single connection or a pool. Internally, _internal.get_connection does two different things depending on which you passed. Given a ConnectionPool, it checks a connection out for each operation and returns it afterwards. Given a Connection, it yields the same connection forever. So a saver built by the factory, or on a connection you checked out of a pool and handed over, holds one socket across every LLM call, tool call and human wait for as long as the process lives. A saver built on a pool holds a socket only while it reads or writes.
Two constraints come with the constructor route. The README requires autocommit=True and row_factory=dict_row on any connection you pass in. Autocommit is needed because setup() applies 10 migrations, three of which are CREATE INDEX CONCURRENTLY statements that Postgres refuses inside a transaction block. dict_row is needed because the saver reads columns by name. And passing a pool together with a pipe raises ValueError("Pipeline should be used only with a single Connection, not ConnectionPool.").
One more property sets expectations for pool sizing. _cursor holds self.lock (a threading.Lock in the sync saver, an asyncio.Lock in the async one) for the whole time it holds a connection, and every checkpoint read and write goes through _cursor. It follows from the source that one saver instance uses at most one pooled connection at a time. Raising max_size adds no checkpoint throughput for a single saver. The pool is there for validation, recycling and reconnect, and for sharing with an AsyncPostgresStore if you run one.
Store vs checkpointer
PostgresStore holds cross-thread memory; the checkpointer holds per-thread state. They ship in the same package and differ in exactly the two places this post cares about. PostgresStore.from_conn_string accepts pool_config, builds a ConnectionPool with prepare_threshold 0 by default but merges your kwargs last, so you can override it there, and supports TTLConfig with default_ttl and sweep_interval_minutes. The checkpoint savers have neither. If you run both, an AsyncPostgresStore can share the async pool built further down. MemorySaver keeps state in process memory and loses it on restart: tests only.
LangGraph Postgres checkpointer errors, mapped to causes
These are the six symptoms that trace back to the 3.1.2 code path, with the mechanism, the fix you control, and where upstream stands as of 2026-10-05.
The first, fifth and sixth rows have deterministic fixes. The second has a reliable mitigation. The third and fourth are open bugs, covered further down. Every fix is client configuration, which matters where you cannot change the pooler or the firewall.
| Symptom | Mechanism at 3.1.2 | Fix on your side | Upstream state |
|---|---|---|---|
| Prepared-statement errors through PgBouncer or another transaction-mode pooler | Factory hardcodes prepare_threshold=0, so every statement is prepared on first use, and the next transaction can land on a server session that never saw it | prepare_threshold=None on the connection or pool, or the PgBouncer 1.22+ path | No issue needed; the factory has no argument for it |
the connection is closed after the agent sits idle | One connection held for the process lifetime; a firewall, load balancer or proxy drops it while idle | A pool with check and max_lifetime; libpq keepalives | #7304 open, PR #7311 closed unmerged |
SSL error: bad length plus SSL SYSCALL error: EOF detected | Not diagnosed upstream | Validating pool and keepalives narrow the window; no confirmed fix | #3716 open, 53 comments |
SSL connection has been closed unexpectedly with AsyncPipeline [BAD] in the logs | Not diagnosed upstream; writes still pipelined despite pipeline=False | Validating pool, supports_pipeline = False, retry on connection errors | #5675 open, PR #8020 closed unmerged |
connection is closed on writes behind PgBouncer in transaction mode | supports_pipeline comes from client libpq 14+, so writes enter pipeline mode regardless of the pooler | checkpointer.supports_pipeline = False after construction | #8420 open, PR #8421 closed unmerged |
TypeError: tuple indices must be integers or slices, not str | Your own connection or pool lacks row_factory=dict_row | Add dict_row (and autocommit=True) to the connection kwargs | Documented in the README |
PgBouncer and prepared statements: prepare_threshold=None
psycopg's semantics are precise, and the factory picks the most aggressive value. From the prepare_threshold docstring in psycopg 3.3.6: it is the number of times a query is executed before it is prepared; at 0, every query is prepared the first time it is executed; at None, prepared statements are disabled on the connection; the default is 5. So from_conn_string asks Postgres to prepare every checkpoint statement on first use.
That is fine against a direct connection. Behind a transaction-mode pooler it is the wrong setting. The psycopg documentation on prepared statements says that unless a connection pooling middleware explicitly declares otherwise, it is not compatible with prepared statements, "because the same client connection may change the server session it refers to", and that with such middleware you should set prepare_threshold to None. The warning is generic, so it covers RDS Proxy and Supavisor as much as PgBouncer unless their own documentation says otherwise.
There is a path that keeps prepared statements. Starting from psycopg 3.2, they work through PgBouncer when all three conditions hold:
max_prepared_statements is greater than 0.Capabilities.has_send_close_prepared() reports.With PgBouncer 1.22+ but an older libpq, psycopg documents a fallback: disable deallocation by setting Connection.prepared_max to None (its default is 100 statements per connection). Start from prepare_threshold=None, and move to the 1.22 path only after the platform team confirms the pooler config and you have pinned libpq in your image.
If you cannot replace the factory this week, two attributes can be changed after construction. saver.conn is the psycopg connection and prepare_threshold has a setter; the reporter of #5675 uses exactly this line. supports_pipeline is covered in the next section.
async with AsyncPostgresSaver.from_conn_string(DB_URI) as checkpointer:
checkpointer.conn.prepare_threshold = None # factory hardcodes 0
checkpointer.supports_pipeline = False # pipeline=False alone does not do this
graph = builder.compile(checkpointer=checkpointer)This patch fixes the prepared-statement row and the pipeline row of the table. It does nothing for dead idle connections, because the saver still holds one connection for the process lifetime. Treat it as a stopgap until the pool construction below replaces it.
Pipeline mode: why pipeline=False does not turn it off
Both constructors run self.supports_pipeline = Capabilities().has_pipeline(). The same line appears in the sync saver, the async saver and both Shallow savers. has_pipeline() is a psycopg capability check on the client: to use pipeline mode, "the client must use a libpq from PostgreSQL 14 or higher". It knows nothing about the server, PgBouncer, RDS Proxy or the network in between. Issue #8420's description says it checks the server; it does not.
Then look at _cursor(pipeline=...), identical in sync and async:
pipe, use it and sync after pipelined calls.pipeline=True and supports_pipeline is true, wrap the cursor in conn.pipeline().pipeline=True and supports_pipeline is false, wrap it in conn.transaction().put, put_writes and delete_thread (and their async versions) call _cursor(pipeline=True). Reads (get_tuple, list, setup) use a plain cursor. So on any libpq 14 or newer, every checkpoint write runs in pipeline mode: when you passed pipeline=False, when you passed nothing, and when conn is a pool. The factory's pipeline flag only decides whether the whole saver lives inside one long-lived pipeline.
psycopg's own documentation calls pipeline mode experimental: "Its behaviour, especially around error conditions and concurrency, hasn't been explored as much as the normal request-response messages pattern." #8420 reports connection is closed behind PgBouncer in transaction mode and gives supports_pipeline = False as the workaround. Pipelining is an extra moving part on a path that already has a pooler and a firewall in it, and you can remove it.
The workaround is one line after construction. supports_pipeline is a plain public instance attribute read on every _cursor call, so setting it to False switches writes to conn.transaction(). #8420 proposed a constructor parameter; its PR, #8421, was closed without merge on the day it was opened, 2026-07-23. The sync PostgresSaver is affected in exactly the same way. #8420 claims the sync saver uses psycopg2 and is immune; at 3.1.2 it uses psycopg 3 and carries the identical line.
Connection is closed after idle: build your own pool
A run waits on a slow model, a long tool call or a human approval, and a single connection held across that wait is exposed to every idle timer on the path: the pooler, a load balancer, a firewall. Ask your network team for the idle timeout on each hop before you pick numbers; provider defaults pasted in issue comments are not a substitute.
This construction replaces from_conn_string for the async saver. Each line maps to a requirement covered above.
from psycopg.rows import dict_row
from psycopg_pool import AsyncConnectionPool
from langgraph.checkpoint.postgres.aio import AsyncPostgresSaver
# sslmode=verify-full and the CA path are your PKI's, not defaults
DB_URI = "postgresql://agent_rw@pgbouncer.internal:6432/agents?sslmode=verify-full"
pool = AsyncConnectionPool(
conninfo=DB_URI,
min_size=1,
max_size=4,
max_lifetime=1800, # bounds connection age; illustrative, below psycopg's 3600s default
check=AsyncConnectionPool.check_connection, # validate on every checkout
kwargs={
"autocommit": True,
"row_factory": dict_row,
"prepare_threshold": None, # transaction-mode pooler: no prepared statements
},
open=False,
)
async def build_checkpointer() -> AsyncPostgresSaver:
await pool.open()
checkpointer = AsyncPostgresSaver(pool)
# Behind PgBouncer/RDS Proxy/Supavisor: route writes through conn.transaction()
checkpointer.supports_pipeline = False
return checkpointerConstruct the saver inside a running event loop: AsyncPostgresSaver.__init__ calls asyncio.get_running_loop(). Pass the pool, never a connection you checked out of it for the whole run, or you are back to one long-lived socket.
max_idle=300 appears in the #3716 and #7304 threads as the cure for a reaper timeout. Read the definition: it only shrinks connections that the pool grew beyond min_size. Since one saver uses one connection at a time, the pool usually sits at min_size, which is exactly where max_idle has no effect. check and max_lifetime act on every connection. A check validates at checkout, so a connection can still die between checkout and the write; it narrows the window rather than closing it.
What each pool setting does, and the one that does not help
The psycopg-pool 3.3 documentation defines the parameters, and one widely pasted fix does not do what it is pasted for.
libpq keepalives
The other lever sits below psycopg, in libpq's connection parameters, which you can put in the connection string. keepalives defaults to 1 (on). keepalives_idle, keepalives_interval and keepalives_count control when the OS starts probing an idle socket, how often, and how many failures end it; for each, a value of zero uses the system default, and all of them are ignored over Unix-domain sockets. tcp_user_timeout, in milliseconds, bounds how long transmitted data can stay unacknowledged. The PostgreSQL documentation gives no recommended values, and neither will we. The rule is relational: the idle value has to be shorter than the shortest idle timeout on the path, or the probe arrives after the connection is already gone. Support is platform-dependent, so check the libpq page for your OS.
Retries are safe for the checkpoint, not for the node
Once a write fails, retrying it is tempting. For the checkpoint rows it is safe: the upserts at 3.1.2 use ON CONFLICT ... DO UPDATE for checkpoints, DO NOTHING for blobs, and one of the two for writes, so a retried write with the same ids creates no duplicate rows. That says nothing about the node that produced the state. If the node sent an email or posted a payment instruction before the checkpoint write failed, a replay sends it again. Design those side effects as described in retry, reroute and replan for agent tool failures; the checkpointer is LangGraph's durability layer, and the general theory of replay and idempotency is in durable execution for AI agents.
| Setting | Default | What it does for a checkpointer |
|---|---|---|
check | No check | Calls the callback on every getconn() or connection(); check_connection (added in 3.2) discards a dead connection before the saver uses it |
max_lifetime | 1 hour, reduced at random by up to 5% | Closes and replaces connections past a fixed age, so none lives indefinitely |
max_idle | 10 minutes | Closes idle connections only above min_size, so it does nothing for a pool sitting at min_size |
reconnect_timeout | 5 minutes | How long the pool keeps trying to reconnect before calling reconnect_failed |
SSL error bad length and SSL connection closed unexpectedly: still open
Two of the six rows have no root cause on record, so what the evidence rules out is the useful part.
#3716: SSL error: bad length
Opened 2025-03-06, open on 2026-10-05, 53 comments and 12 upvotes, last updated 2026-08-17. The full error is psycopg.OperationalError: sending query and params failed: SSL error: bad length, often with SSL SYSCALL error: EOF detected. The only contributor replies, on the day it was opened, asked about disk space and large tool messages. No maintainer has posted a diagnosis or a fix since. What the thread rules out: What the thread does show: the reporters who name their database are on managed cloud Postgres (Azure, AWS RDS, Cloud SQL), one found it occurred on cloud databases but not on local ones, and the recurring precursors in the logs are error ignored terminating <psycopg.AsyncPipeline [BAD]>: the connection is lost and psycopg.pool | discarding closed connection. The leading community explanation is a connection killed underneath a pipelined write. It is plausible and unconfirmed. One commenter, in August 2026, reports reproducing the error locally on 3.1.2 and psycopg 3.3.4 by terminating the application's backends with pg_terminate_backend from a second session between checkpoint writes, and that a pool with check=AsyncConnectionPool.check_connection recovered. You can run the same procedure on a staging database: start a long run, kill its backend from psql mid-run, and watch whether the next write fails or reconnects.
- Prepared statements. The original reporter already built an
AsyncConnectionPoolwithautocommitandprepare_threshold0, across versions 2.0.9 to 2.0.15. One commenter reports thatprepare_threshold=Nonedid not help. That fits the mechanism: prepared statements are a session setting, not a connectivity setting. Do not expectNoneto fix this error. - A single version. It spans several releases and has been reported on 3.1.2.
#5675: SSL connection has been closed unexpectedly
Opened 2025-07-26, open, 17 comments, last updated 2026-06-08. The error is psycopg.OperationalError: consuming input failed: SSL connection has been closed unexpectedly, alongside psycopg.AsyncPipeline [BAD]. The reporter connects through the Supabase pooler on port 6543, passes pipeline=False, sets checkpointer.conn.prepare_threshold = None, and still fails with minimal state. The source narrows what this rules out. pipeline=False did not stop the reporter's writes from being pipelined; supports_pipeline was still true. So #5675 does not show that the failure happens without pipeline mode, only that it happens without a saver-level pipe and without prepared statements. The reporter was also on the single-connection factory, with no validation on use. One commenter saw the error stop after shrinking large state; the reporter said that did not apply to them. Large state is worth cutting anyway, since every write carries it; context compaction for long-running agents covers offloading big payloads out of graph state. Later comments converge on a validating pool plus a saver subclass that retries on connection errors. PR #8020 proposed disabling autocommit under pipelines and was closed without merge on 2026-06-07; the shipping factory still uses autocommit with pipelining.
What to log while these stay open
Raise the psycopg and psycopg.pool loggers to WARNING and alert on the two precursor lines above. A checkpoint write that fails and is swallowed leaves a run that looks finished but cannot be resumed or audited, which is the failure class covered in detecting silent AI agent failures. Plot write failures against the pooler and firewall change log.
Deleting old checkpoints: retention in regulated deployments
Here the on-prem context changes the answer most. In a bank, insurer or hospital, the agent's state lands in the institution's own Postgres, and the checkpoint tables are a step-level record of what the agent saw and did: model inputs, tool arguments, tool outputs, every intermediate state. That is a records object with an owner, an access policy and a retention schedule, and at 3.1.2 the package gives you none of them by default.
The safe unit is the whole thread. Select threads whose newest root checkpoint is older than your retention window, then delete each through the saver so all three tables stay consistent:
STALE_THREADS_SQL = """
SELECT thread_id
FROM checkpoints
WHERE checkpoint_ns = ''
GROUP BY thread_id
HAVING max((checkpoint->>'ts')::timestamptz) < now() - %s::interval
"""
async def sweep(pool, checkpointer, retention: str = "90 days") -> int:
async with pool.connection() as conn:
rows = await (await conn.execute(STALE_THREADS_SQL, (retention,))).fetchall()
for row in rows:
await checkpointer.adelete_thread(row["thread_id"])
return len(rows)The "90 days" is a placeholder; the value comes from your records schedule, not from us. The ts expression is unindexed, so run the sweep off-peak or add an expression index. If audit requires the history to outlive the operational window, export threads to the records system before the sweep deletes them. It is the same shape of problem as an LLM gateway's spend logs table: an append-only Postgres table with no default retention, where the obligation should set the window and disk pressure should not.
The same review should settle three more decisions.
sync, async (the default) and exit. exit persists only when the graph exits, which writes far fewer rows and loses the step-level trail. ShallowPostgresSaver still ships in 3.1.2, with a warning that it is deprecated in favour of PostgresSaver with durability="exit".setup() creates tables and indexes and records versions in checkpoint_migrations. LangGraph's docs recommend running migrations as a dedicated deployment step; in a change-controlled estate, run it with a separate DDL role so the agent's runtime role only reads and writes.LANGGRAPH_STRICT_MSGPACK=true or an explicit allowed_msgpack_modules list to restrict checkpoint deserialization to known-safe types; the SQLite saver's CVE is covered in the LangChain and LangGraph CVE audit.For how checkpointing fits into the rest of an agent architecture, see the AI agents pillar.
| Question | Answer at 3.1.2 |
|---|---|
| Is there a TTL? | No. TTLConfig exists on PostgresStore only |
| Is there a timestamp column? | No. The time is checkpoint->>'ts' inside the JSONB, ISO 8601, unindexed |
| What can delete? | delete_thread / adelete_thread: whole thread, all namespaces, three tables |
What about prune, delete_for_runs, copy_thread? | Defined on the base class, not overridden, raise NotImplementedError |
| Can I delete old rows inside a live thread with SQL? | Unsafe: the base class warns that a naive prune can leave DeltaChannel channels to "silently reconstruct as empty", with no error raised |
What to do this week
Grep your services for from_conn_string. For each hit, replace it with the pool construction: autocommit=True, row_factory=dict_row and prepare_threshold=None in the kwargs, check=AsyncConnectionPool.check_connection, a max_lifetime you chose deliberately, and supports_pipeline = False on the saver if any pooler sits between the agent and Postgres. Ask the network team for the idle timeout on every hop and set the libpq keepalive idle value below the shortest one. Then run the pg_terminate_backend drill on staging and confirm the next write recovers. Last, take the stale-thread query to whoever owns the records schedule, agree the window, and schedule the sweep.
FAQ
Quick answers to the questions this post tends to raise.



