In our last BLOG about IBM acquiring Confluent we tried to paint the picture of the ‘Background’ and capabilities at high level of Confluent and Apache Kafka and if you did not read it then perhaps a 2-3 minute reviewing what we reflected that might be worthwhile. See here;
Why IBM Spent $11BN on Acquiring ‘Confluent’ – SmallNet
In this short post we wanted to try and position how and where confluent fits in on day one at IBM which is extremely complimentary and most likely is used by many of you today, perhaps not as the Confluent platform but Apache Kafka the messaging system.
Its estimated that there are over 150,000 sites/clients using the open source solution ‘Apache Kafka’ and by Confluent ‘productionising Kafka into a fully supported platform its believed that this data streaming platform that more than 6,500 enterprises, including 40% of the Fortune 500, rely on to power real-time operations.
What has made Confluent even more appealing of late is the development of ‘Kora’ which is not an Apache project itself, but rather a cloud-native engine developed by Confluent to power Confluent Cloud’s fully managed Apache Kafka service.
It completely redesigns the traditional Kafka architecture to deliver elastic scalability, high performance, and multi-tenancy across major cloud providers.
Key Features of ‘Kora’ are;
- Elastic Scalability: Scales up to 30 times faster than open-source Apache Kafka.
- Low Latency: Delivers tail latencies up to 10–16x faster than standard open-source Kafka deployments.
- Cloud Efficiency: Operates efficiently across multi-region cloud infrastructures like AWS, GCP, and Azure.
- Cost Reduction: Lowers the total cost of ownership (TCO) by up to 60% compared to self-managed setups.
The KEY value proposition to this acquisition is that together IBM and Confluent deliver an extremely smart data platform that gives every AI model, agent, and automated workflow the real-time, trusted data needed to operate across on-premises and hybrid cloud environments at scale. It is also extremely secure to its very core.
MARKET Driven
So, what’s driven this market is that as enterprises are now moving from AI experimentation to production usage, their critical barrier to that success is/was the data — it had to be clean, governed, continuously refreshed —and delivered at the speed and scale that AI demands. That’s exactly what this solution provides……
However, the reality in most enterprises today’s that the data remains siloed across systems and environments, arriving hours or days after it is generated. This combined solution of IBM and Confluent provides the fabric through which AI agents can access the information they need, with all the controls, governance, and real-time velocity to put that information to work safely and at scale.
IT should also be NOTED that if you dig deep into your business, many of your estates will most likely be using ‘open source’ software for various tasks and more than likely Apache Kafka for streaming of messages.
Indeed, we at Smallnet Consulting have several clients using this within their IBM DataStage and CDC (Data Replication) estates which is a very common deployment. This is probably also true of many other ETL/ELT providers in the marketplace.
The challenge is often NOT the actual technology or capability of such ‘open source’ software it’s the DAY to DAY support and maintenance of this, plus building in all the security, governance and the ability to deploy and monitor at scale.
To give you some context to IBM and Confluent deployments today here are a few examples across various differing industries;
- Michelin relies on Confluent to manage real-time inventory across a supply chain spanning 170 countries — achieving 35% cost savings without sacrificing visibility or control.
- L’Oréal uses Confluent to stream real-time product and inventory updates across internal systems and third-party applications, helping the company respond faster to changing consumer demand.
- BMW Group streams IoT data from 30+ production sites and its global sales network in real time, connecting factory floor systems and cloud applications across the organisation.
- Ticketmaster streams ticket inventory, sales, and customer activity in real time across hundreds of systems, reducing development friction and powering machine learning at scale.
As referenced above whilst IBM has acquired Confluent it is used VERY widely across lots of technology platforms; many of these will be with names you know or perhaps even utilise today, such as;
- Hyperscalers like AWS, Azure, Google Cloud and more
- Applications such as SAP, DataBricks, Snowflake, Oracle, Mongo DB, Microsoft Applications,
- Business Intelligence solutions (Qlik/Cognos) and many more particularly when ‘real time data reporting’ is essential.
SO WHERE does Confluent currently FIT within the IBM portfolio on day one and indeed has done on some platforms for many years?
As a data integration SME in this marketplace, we can advise that several of our existing IBM DataStage/CDC (Replication) clients have used Apache Kafka, and the interfaces are well known to that community and supported by great documentation.
We would be delighted to talk with you or your team about the various ‘USE cases’ and indeed there are some great web sites showing these no matter what platform you are running/operating. However, as a ‘snapshot’ of where it can be deployed already within IBM technology exists at the likes of;
IBM Watsonx.data integration (THIS IS A NEW naming convention of a unified approach but includes DATASTAGE) – So this is a new a unified control plane platform designed to consolidate enterprise data tool stacks ie single product for ALL use cases. It manages batch ETL/ELT, real-time streaming, data replication, and unstructured data pipelines while providing native data observability from a single interface.
Core Capabilities include;
- Agentic Integration: Uses artificial intelligence to translate natural language prompts into working pipeline plans and automated executions.
- Unstructured Data Processing: Handles ingestion, PII masking, chunking, embedding, and vectorisation for generative AI readiness.
- Multi-Style Execution: Combines traditional batch processing, real-time streaming, and log-based change data capture (CDC) replication.
- Data Observability: Tracks end-to-end pipeline health, triggers alerts, and provides diagnostic tools to catch failures quickly.
- Architectural Highlights include authoring Flexibility: Supporting low-code/no-code graphical canvases, Python SDKs, SQL interfaces, and AI-assisted creation.
- Hybrid Connectivity: Operates across multi-cloud, on-premises, and SaaS environments without forcing unnecessary data movement
IBM Watsonx.data – This is one IBM latest open platform which is a hybrid, open data store built on a lakehouse architecture. It lets companies store, manage, and analyse both structured and unstructured data for AI and business analytics. It combines the low cost of a data lake with the speed of a data warehouse so is ideally suited to Confluent.
Key Features include;
- Open Architecture: Uses open formats like Apache Iceberg and separates compute from storage to stop vendor lock-in.
- Utilises Multiple Engines: Uses different query tools like Presto and Spark so you can run workloads on the best tool for the job.
- Built-in Governance: Connects security and access policies directly to data sources.
- Hybrid Cloud: Connects to data where it already lives across multi-cloud or on-premises systems.
IBM Watsonx data intelligence – The IBM watsonx.data intelligence platform has evolved from extremely powerful catalog and data quality solutions and now blended/integrated with the Manta Lineage solution which IBM also acquired. This solution adds business context to fragmented enterprise data, making it more usable, much easier to find, trustworthy, and valuable for both AI and business analysts. By connecting metadata, definitions, lineage, quality, and governance, it helps organisations understand, trust, control, and scale data to support confident decisions and AI outcomes.
IBM webMethods (Integration) – IBM webMethods is a unified, hybrid integration platform (iPaaS) used to connect applications, APIs, events, data, B2B/EDI transactions, and file transfers across cloud and on-premises environments and by definition it will use the likes of Kafka. It helps enterprises eliminate data silos, manage complex workflows, and operationalise AI-driven processes again from a single control plane.
IBM MQ – IBM MQ has existed within IBM for many years and is reliably used in many highly regulated industries where ‘trust’ is critical. It is robust message-oriented middleware that lets independent software applications send and receive data securely across different computer systems. It guarantees reliable, once-and-only-once message delivery so that data is never lost, even if a server or application goes offline
IBM Z – The most critical business transactions in the world have long run on IBM Z. With IBM Z and Confluent, organisations can identify and drive real-time events at the transaction source as well as stream transactional data directly for real-time analytics, automation, and AI workflows. This enables mission-critical transaction processing to integrate tightly with the rest of the business in real-time, at enterprise scale.
In summary, the acquisition of Confluent was a natural progression by IBM to ensure ‘messaging data’ in real time across multiple platforms and environments was adding real value to IBM’s clients usage of Kafka. But, also ensuring the security and productionising the ability to deploy across multiple applications.
We have tried to show the history and background to the Confluent offering and importantly the synergy and fit within IBM already with no doubt much more to follow around IBM’s AI and integration initiatives.
If you would like to know more about Confluent’s solutions and/or USE case scenarios or simply some consultancy skills/advise then simply contact us at: get in touch.