Guide
Premium
Intermediate
Real Projects

Exactly-Once in Kafka: The Idempotent Producer, Transactions, read_committed, and the Boundary Everyone Crosses

A practical guide to the guarantee almost everyone thinks they have and very few actually do: the difference between delivering a message once and applying its effect once, the two duplicates that are not the same problem (producer-side retries and consumer-side rebalances), the idempotent producer from the inside with PID, epoch and per-partition sequence numbers — plus the three settings it forces, including the five-in-flight limit and why it exists — transactions with a stable transactional.id, zombie fencing in initTransactions, and sendOffsetsToTransaction as the piece that makes the guarantee real, the complete consume-transform-produce loop in Java and in Python with confluent-kafka, the three details nearly every implementation gets wrong (the seek after an abort, which is silent data loss; the +1 on the offset; and the groupMetadata overload that enables EOS v2), read_committed and the LSO with the offsets that appear to go missing and the aborted messages filtered client-side, Kafka Streams with exactly_once_v2 and the commit.interval.ms almost nobody touches, KIP-890 and transaction.version=2 in Kafka 4.0 with the per-commit epoch bump that closes the cross-transaction message leak, and the boundary that voids all of it the moment you call a payment gateway. With eight recurring mistakes, the hanging-transaction runbook using kafka-transactions.sh, the alert that actually works (the log-end-offset to LSO gap, not lag), the real cost in throughput and latency, a production checklist, FAQ and glossary. Production-ready code in Java, Python and shell. September 20, 2026 expansion: the internals with PID, epoch and per-partition sequence numbers, why the in-flight limit is five, and the UNKNOWN_PRODUCER_ID that looks alarming without being a problem; the transaction coordinator and the __transaction_state log step by step, with PrepareCommit as the point of no return and why a hanging transaction blocks consumers that are not yours; the transactional.id as the decision almost nobody gets right, with the StatefulSet that provides stable identifiers and a closing answer to the historical KIP-447 and groupMetadata() question; what read_committed does to your latency, covering the LSO, transaction size and the aborted-transaction index you pay for in bandwidth; Kafka Connect with exactly.once.source.support and its two-phase rolling restart, and why there is no equivalent switch for sink connectors; Flink with transactional-id-prefix and the timing trap that causes silent data loss; the three patterns for when the guarantee leaves Kafka (inbox, outbox and idempotent writes) with ready-to-use SQL; multi-cluster and why MirrorMaker 2 does not carry the guarantee across; five metrics and two Prometheus alerts; fault-injection tests with Testcontainers and the difference between fatal and retryable errors; a complete case where latency rises 29% and incidents drop to zero; when not to use exactly-once; how to enable transaction.version=2 without surprises; and four questions that always come up afterwards.

37 minutes read
Josué Puig
1 views

Verificando acceso...

Loading comments...

Related Resources

Guía
PREMIUM

API Versioning and Contract Testing: Safe Changes, OpenAPI in CI, Pact, and Retiring Endpoints with Sunset

A practical guide to changing your APIs without breaking the people who consume them: why compatibility rules invert between the request and the response —and why widening a returned enum breaks clients even though you are "only adding"—, the tolerant reader pattern in Pydantic with an escape hatch and a metric, OpenAPI generated from the code and diffed on every pull request with oasdiff (including the git diff --exit-code step without which the whole check is theatre), consumer-driven contract testing with Pact: type matchers instead of literal values, well-designed provider states, version selection with deployed_or_released, and the gate that actually makes it safe, can-i-deploy paired with record-deployment. It also covers rolling this out without stopping the factory using pending and WIP pacts, bi-directional contracts when the provider is a third party, the three versioning strategies with their real operational costs, why versioning the whole API for a single endpoint guarantees nobody migrates, retiring versions with the Deprecation (RFC 9745) and Sunset (RFC 8594) headers plus the migration Link, the per-consumer metric without which no sunset date is ever met, brownouts returning 410 Gone before the final shutdown, and the BACKWARD, FORWARD and FULL compatibility modes for event schemas. With production-ready code in Python, YAML, Bash and PromQL, eight recurring mistakes, a production checklist, FAQ and glossary. It also extends the contract beyond the happy path: errors with RFC 9457 (problem+json), cursor pagination and defaults as part of the contract, the expand/contract pattern for renaming a field across database and API with no maintenance window, governance with Spectral, buf breaking for gRPC and Protobuf, GraphQL schema evolution with @deprecated and real per-field usage, and semantic versioning of generated SDKs. It also covers the contracts that never show up in the schema and break just as hard: outbound webhooks with the version pinned on the subscription, HMAC signing with a timestamp window and key rotation, Idempotency-Key for safe POST retries, the RateLimit and RateLimit-Policy headers —and why lowering a limit is a breaking change—, and OAuth scopes, the blind spot no OpenAPI diff will ever catch.

Guía
PREMIUM

asyncio in Production: Never Block the Event Loop — TaskGroups, Cancellation and Bounded Concurrency

A practical asyncio guide for Python services in production: why blocking the event loop degrades the whole process without raising a single exception, how to catch it by measuring loop lag and with Python 3.14 introspection, structured concurrency with TaskGroup and handling ExceptionGroup via except*, the task the garbage collector makes vanish, timeouts with a deadline budget propagated across services, correct cancellation with cleanup and shield, bounded concurrency with semaphores and backpressured queues, synchronization primitives, and what changes with eager tasks, python -m asyncio pstree and free-threading. With production-ready code and a deployment checklist.

Guía
PREMIUM

Cache-Aside in Production: TTLs, Invalidation, and How to Prevent Cache Stampedes

The complete guide to the cache-aside pattern with Redis: jittered TTLs, correct invalidation, and the three defenses against cache stampedes (distributed lock, single-flight, and XFetch). Expanded with stale-while-revalidate, fail-open and circuit breakers, two-tier caching with RESP3 invalidation, delayed double delete and CDC, hot keys, eviction and memory management, observability with Prometheus, testing, choosing an engine (Redis, Valkey, Memcached), and a complete TypeScript implementation. With production-ready code in Python and TypeScript.