Tutorials

Apache Kafka Tutorial

If you're curious about how Kafka stacks up against other options, check out the JMS vs Kafka article to better understand the unique advantages Kafka offers.

Kafka is a powerful messaging system used for streaming and processing big data. It's fast, scalable, and fault-tolerant, making it a preferred choice for companies such as PayPal and Uber. This tutorial introduces you to Kafka's basic architecture and demonstrates publishing and consuming messages.

Key Takeaways

  • Kafka is a scalable and fault-tolerant messaging system used for data streaming.
  • Its architecture features producers, consumers, topics, partitions, and brokers.
  • Partitions enable Kafka to handle parallel processing efficiently.
  • This tutorial guides you through installing Kafka, and sending and consuming messages.

What is Kafka?

Kafka is a messaging system that securely transfers data from one system to another. It operates on a distributed cluster of servers, ensuring high availability and reliability for data streaming.

Kafka Architecture

Producers send data to Kafka topics, while consumers retrieve data from these topics. Kafka topics are divided into multiple partitions across different brokers or servers in the cluster. Each partition is replicated across brokers to safeguard against data loss. Consumers read data from individual partitions, enabling parallel processing. Here's a breakdown of Kafka's key components:

Kafka Cluster

A Kafka cluster is the collection of servers on which Kafka operates.

Topic

Topics categorize data in Kafka. You write to and read from specific topics in Kafka.

Partition

Data in topics is distributed over partitions. Partitions organize records sequentially from oldest to newest, forming the topic structure.

Broker

Brokers are the servers or nodes within the cluster. Partitions are distributed among brokers to ensure load balance.

Producer

Producers send data to topics in Kafka.

Consumer

Consumers read data from topics across the brokers.

Replica

Partitions are replicated across brokers for fault tolerance. This replication is achieved through replicas.

Leader

Among replicated partitions, one is designated the leader. The leader handles all read/write operations for its replicas across brokers.

Follower

Followers are non-lead replicated partitions. They take over for leaders upon failure and otherwise mirror the leader's data.

When a producer writes messages to a Kafka topic, the messages get distributed across the topic's partitions evenly. For example, three incoming messages for a topic with three partitions will each be stored in a distinct partition.

Partitions are replicated within the cluster, with some designated as leaders and others as followers. Producers write to leaders, facilitating effective load distribution across brokers. Kafka aims to balance leaders throughout the cluster to enhance efficiency.

In the event of broker failure, a follower partition becomes the leader to prevent data loss. By distributing partitions across servers, Kafka allows topics and consumers to be processed in parallel. Within consumer groups, distinct consumers can read from unique partitions, enabling entire topic consumption.

Kafka Tutorial

Installing Kafka

Before installing Kafka, ensure Java is set up. To download Kafka, go to the official Kafka downloads page to get the latest version.

tar -xzf kafka_version.tgz
cd kafka_version

This extracts Kafka and its dependencies, including Apache ZooKeeper.

Next, start ZooKeeper, which Kafka relies on for configuration:

bin/zookeeper-server-start.sh config/zookeeper.properties

This starts a local ZooKeeper server.

Creating a Kafka Topic

With ZooKeeper running, execute this command in another terminal:

bin/kafka-topics.sh --create --bootstrap-server localhost:9092 --replication-factor 1 --partitions 1 --topic sample

It creates a topic called sample. You can specify parameters like partitions and replication-factor as needed.

Kafka Producer: Sending a Message

Kafka includes a producer script for sending messages, where each line represents a new message:

bin/kafka-console-producer.sh --broker-list localhost:9092 --topic sample
Hello Kafka!

This starts a producer in the terminal. Specify the topic and type Hello Kafka!, then press enter to send.

Kafka Consumer: Reading the Message

To consume messages sent by the producer, start a Kafka consumer in a separate terminal:

bin/kafka-console-consumer.sh --bootstrap-server localhost:9092 --topic sample --from-beginning
Hello Kafka!

This prints the producer's message to the console. Enter additional messages in the producer terminal to see them display in the consumer terminal.

Conclusion

Kafka uses a server cluster to distribute messages across nodes. This setup provides fault-tolerance and maximizes throughput. By following this guide, you've successfully used Kafka to produce and consume messages!

FAQ

What's the difference between a leader and a follower in Kafka?

A leader partition handles all read and write operations for its data and is responsible for synchronizing its followers. Followers replicate data from leaders and take over if a leader fails.

Why does Kafka require ZooKeeper?

ZooKeeper manages configuration, cluster metadata, and controller election in Kafka. Recent Kafka versions are working towards removing this dependency, introducing its native Kafka Raft Metadata mode.

Can Kafka be used without Java?

No, Java is necessary for running Kafka as it is built on Java Virtual Machine (JVM).

Is a single broker sufficient for a Kafka setup?

A single broker setup works for development or testing but isn't ideal for production due to its lack of redundancy and fault tolerance.

Mastering the tech interviewWhat everyone is doing wrong in tech interviews