corpus.blog
Most cited
Talked about
Blogs
41,554 blogs · 4,904,241 posts
Claim your blog
Back to Netdata Blog
Blog · corpus.blog/blogs/netdata.cloud/posts
Netdata Blog
netdata.cloud
Undated
Kafka Broker May Not Be Available: How To Fix It
original ↗
—
Kafka Broker Out Of Disk: How To Fix It
original ↗
—
Kafka CommitFailedException: rebalanced-out consumers and poll loop timeouts
original ↗
—
Kafka consumer group lag growing: detection, lag-as-time, and root causes
original ↗
—
Kafka consumer group rebalancing too often: heartbeats, session timeout, and assignors
original ↗
—
Kafka Consumer Lag Monitoring
original ↗
—
Kafka consumer rebalance storm: stuck in PreparingRebalance and max.poll.interval.ms
original ↗
—
Kafka controller event queue backing up: overwhelmed controller and stalled metadata
original ↗
—
Kafka disk I/O latency high: await, LocalTimeMs, and the slow-disk broker
original ↗
—
Kafka FailedProduceRequestsPerSec rising: the single best 'producers are hurting' signal
original ↗
—
Kafka fetch request latency high: FetchConsumer vs FetchFollower and page cache misses
original ↗
—
Kafka ISR shrinking: IsrShrinksPerSec, flapping, and the cascade to offline
original ↗
—
Kafka JVM heap and Full GC pauses: ISR drops, session timeouts, and right-sizing the heap
original ↗
—
Kafka KRaft quorum has no leader: current-leader = -1 and frozen metadata
original ↗
—
Kafka LEADER_NOT_AVAILABLE: causes during elections, restarts, and topic creation
original ↗
—
Kafka log compaction falling behind: the dead cleaner thread and unbounded disk growth
original ↗
—
Kafka Log directory failed / OfflineLogDirectoryCount > 0: disk errors and JBOD recovery
original ↗
—
Kafka min.insync.replicas and acks: configuring durability you actually have
original ↗
—
Kafka Monitoring
original ↗
—
Kafka monitoring checklist: the signals every production cluster needs
original ↗
—
Kafka NetworkProcessorAvgIdlePercent low: network thread saturation and TLS overhead
original ↗
—
Kafka NOT_LEADER_FOR_PARTITION: stale metadata, controller lag, and client retries
original ↗
—
Kafka NotEnoughReplicasException: acks=all writes rejected below min.insync.replicas
original ↗
—
Kafka OfflinePartitionsCount > 0: partitions with no leader and how to recover
original ↗
—
Kafka OffsetOutOfRangeException: when retention deletes data before the consumer reads it
original ↗
—
Kafka RecordTooLargeException / MESSAGE_TOO_LARGE: message size limits across the path
original ↗
—
Kafka replica MaxLag growing: slow followers and replica fetcher health
original ↗
—
Kafka REQUEST_TIMED_OUT: produce requests that expire before replication completes
original ↗
—
Kafka Too many open files: file descriptor exhaustion from segments and connections
original ↗
—
Kafka UncleanLeaderElectionsPerSec > 0: confirmed silent data loss
original ↗
—
Kafka UnderMinIsrPartitionCount: confirming the write path is blocked
original ↗
—
Kafka UnderReplicatedPartitions > 0: the most important metric and how to clear it
original ↗
—
Kafka ZooKeeper Monitoring
original ↗
—
Kannel Monitoring
original ↗
—
Keepalived Monitoring
original ↗
—
Kernel Same-page Merging (KSM) Monitoring
original ↗
—
Kube-proxy Monitoring
original ↗
—
Kubelet Monitoring
original ↗
—
Kubeproxy Monitoring
original ↗
—
Kubernetes API Server etcd Latency: How To Fix It
original ↗
—
Kubernetes API Server Rate Limited: How To Fix It
original ↗
—
Kubernetes API Server Slow Or Unresponsive: Causes & Fixes
original ↗
—
Kubernetes Cluster State Monitoring
original ↗
—
Kubernetes conntrack exhaustion: dropped connections under load
original ↗
—
Kubernetes Controller-Manager Leader Election Failures
original ↗
—
Kubernetes DNS resolution failures inside pods
original ↗
—
Kubernetes eviction cascade: when one node failure takes down the cluster
original ↗
—
Kubernetes kube-proxy iptables sync stall: causes and recovery
original ↗
—
Kubernetes kube-proxy IPVS: stale rules and session affinity issues
original ↗
—
Kubernetes kubelet certificate expired: detection, rotation, and recovery
original ↗
—
Kubernetes kubelet memory leak: detection and OOM cycle
original ↗
—
Kubernetes kubelet not responding: PLEG, runtime, and certificate issues
original ↗
—
Kubernetes Monitoring Checklist: The Signals Every Production Cluster Needs
original ↗
—
Kubernetes node CPU saturation: load, throttling, and runqueue depth
original ↗
—
Kubernetes node DiskPressure: detection, eviction, and recovery
original ↗
—
Kubernetes node MemoryPressure: detection, eviction order, and prevention
original ↗
—
Kubernetes node NotReady: kubelet, runtime, and network diagnosis
original ↗
—
Kubernetes node PIDPressure: detection and remediation
original ↗
—
Kubernetes PLEG is not healthy: runtime stalls and node degradation
original ↗
—
Kubernetes pod CrashLoopBackOff: causes, diagnosis, and fixes
original ↗
—
Kubernetes pod creation fails: admission, quota, and CRI errors
original ↗
—
Kubernetes pod Evicted: detection, root cause, and prevention
original ↗
—
Kubernetes pod exits immediately: how to diagnose it
original ↗
—
Kubernetes pod ImagePullBackOff: registry, auth, and network diagnosis
original ↗
—
Kubernetes pod OOMKilled: cgroup limits, evictions, and fixes
original ↗
—
Kubernetes pod stuck ContainerCreating: volume, network, and image issues
original ↗
—
Kubernetes pod stuck on volume mount: CSI, permissions, and timeouts
original ↗
—
Kubernetes pod stuck Pending: scheduling failures explained
original ↗
—
Kubernetes pod stuck Terminating: finalizers, grace periods, and force delete
original ↗
—
Kubernetes PVC stuck Pending: storage class, provisioner, and quota
original ↗
—
Kubernetes Scheduler Not Scheduling Pods: Queue Depth & Failure Reasons
original ↗
—
Kubernetes Service Not Reachable: Kube-Proxy, Endpoints & DNS
original ↗
—
License expiry silently disabling features: monitor days-to-expiry
original ↗
—
Lighttpd Monitoring
original ↗
—
Linode Monitoring
original ↗
—
Linux Sensors Monitoring
original ↗
—
Litespeed Monitoring
original ↗
—
Locating endpoints behind NAT and wireless: the positioning problem
original ↗
—
Logstash Monitoring
original ↗
—
Loki Monitoring
original ↗
—
Lustre Metadata Monitoring
original ↗
—
LVM logical volumes Monitoring
original ↗
—
Lynis Audit Reports Monitoring
original ↗
—
MaxScale Monitoring
original ↗
—
MegaCLI MegaRAID Monitoring
original ↗
—
Meilisearch Monitoring
original ↗
—
Mesos Monitoring
original ↗
—
Microbursts: catching sub-second congestion that minute averages hide
original ↗
—
Microsoft Exchange Server Monitoring
original ↗
—
Microsoft SQL Server (MSSQL) Monitoring
original ↗
—
Microsoft SQL Server monitoring checklist: the signals every production instance needs
original ↗
—
Minecraft Monitoring
original ↗
—
MISCONF Redis is configured to save RDB snapshots - what it means and how to fix it
original ↗
—
Modbus Protocol Monitoring
original ↗
—
MogileFS Monitoring
original ↗
—
MongoDB Application Thread Evictions: How To Fix
original ↗
—
MongoDB Balancer Stuck On Jumbo Chunks: Fix It
original ↗
—
MongoDB cache too small: sizing the WiredTiger cache for your working set
original ↗
—
MongoDB checkpoint duration climbing: diagnosing slow WiredTiger checkpoints
original ↗
—
MongoDB checkpoint stall write freeze: when all writes stop with no error
original ↗
—
Previous
Next