The augmented operator: Why the last decades in operations have led to AI

The CEO Forum 2021 (Big Bets 2021)

In payment processing, the stakes are high, and uptime is not a mere metric—it is the foundational requirement for global commerce. Over the last two and a half decades, the role of keeping the lights on has undergone a fundamental transformation. The complexity of modern IT environments has officially scaled beyond human manual capacity. 
  
The industry is entering an era where the machine must augment the human, not simply by providing more data, but by providing actionable clarity. We must take time to explore the critical intersection of observability, AIOps and the human factor—three pillars that can no longer be managed in isolation if operational stability is to be maintained. This timeline traces how it’s all evolved over my career.

The evolution of complexity: From racks to reflexes

To understand the current state of industry leadership in operations, one must analyze the three distinct eras that shaped current challenges. Let’s discuss how observability debt has accumulated and created the need for a new paradigm.  

The era of iron (2000s): Singular management 

In the late 1990s and early 2000s, observability was a physical reality: a walk to the server room and a direct eye on the screen. Hardware was managed in singularity—one box, one operating system (OS), one application. 

At the time, monitoring was binary (up or down). While rigid, it was manageable because the primary constraint was the physical speed of human hands and the number of screens a single operator could monitor. In this era, the human served as the primary integration layer. 

The rise of the black box (2010s): The virtualization debt 

The introduction of virtualization and cloud computing has fundamentally changed the DNA of operations. The shift from physical servers to instances provided incredible agility but created a massive visibility gap. Organizations began moving from simple pings to complex logging and metrics yet remained drowning in data while starving for insights. 

Most teams were monitoring the what—knowing exactly when a CPU spiked—without the tools to understand the why. This era birthed the dashboard fatigue that continues to plague modern ops teams. It proved that more data does not equal better results; it often just leads to faster notification of a disaster. 

The complexity explosion (2020s): The hybrid reality 

Today’s landscape involves microservices, serverless architectures and ephemeral containers that exist for only seconds. Despite the move toward cloud-native builds, data centers remain a critical anchor, especially within regulated industries like global finance. The hybrid path is the current reality: a high friction mix of legacy reliability and modern scalability. 

Traditional monitoring often struggles to surface issues in distributed environments.
In a distributed, non-linear system, observability is required. It is the ability to use internal telemetry to trace a single transaction through a maze of services and identify a single millisecond of latency before it triggers a butterfly effect across the entire ecosystem. 

The strategic target: AIOps and the human factor 

As organizations integrate AIOps (artificial intelligence for IT operations), a common industry anxiety persists: the fear that the human is being removed from the equation. However, historical evidence and the relentless pressures of modern delivery suggest the opposite. 
  
Human capacity is linear, but system complexity is exponential. AI is not a replacement for the operator; it is an essential augmentation. It serves as the filter that allows human expertise to remain effective above a flood of telemetry. 
The evolution of a mature operations department must focus on three specific pillars: 

Pillar one: AIOps as a filter—reclaiming the golden hour

In a standard incident, 10,000 alerts can flood a dashboard in seconds—a phenomenon known as an alert storm. This triggers immediate cognitive overload. In most organizations, the first golden hour of an outage is wasted simply trying to identify the root cause amidst the noise. 

AIOps can help teams consolidate large volumes of alerts, depending on configuration.
Filtering the noise to find the signal reclaims the most precious resource in an incident: time. These filters allow the team to focus on high-level problem solving rather than data sorting. 

Pillar two: The human element and biased observability 

Standard site reliability engineering (SRE) focuses on logs, metrics and traces. While these tell the story of the machine, they are often blind to the human element. The human pillar must be integrated as a formal part of any observability strategy and consist of two core items: 

•    Operational health: If dashboards are green but the engineering team is experiencing high burnout and turnover, the system is fundamentally failing. A sustainable system requires a sustainable team. 
•    The risk of biased human observability: In highly experienced departments, systems often appear more stable than they are. This is due to experts performing invisible fixes—small, manual ticket micro-fixes or adjustments based on intuition before a threshold is reached. 

These human heroics create a dangerous bias that hides true architectural weaknesses. AIOps acts as a corrective lens by observing these repetitive human interventions and surfacing them as patterns. Exposing biased human observability allows an organization to move from heroic manual effort toward systemic architectural health.

Pillar three: Transitioning to augmented operators 

For decades, the badge of honor in IT operations has been the reactive firefighter; the team that fixes emergencies or putting out fires. While culturally celebrated, this model is unsustainable. 

The strategic goal is the transition to augmented operators. This transition involving AIOps usage can help organizations move from reactive operations to more proactive practices and approaches that support predictive insights.

 An augmented operator does not just watch a screen; they orchestrate a system that assists in its own healing, freeing human talent for innovation and architectural advancement. 

The mandate for change 

The stakes in IT operations—particularly within sectors like the IBM Payments Center® that guard the digital economy—have never been higher. Relying on the tools of the previous decade to manage the complexity of the current one is an invitation to failure. 

The industry requires systems that provide context, not just alerts. More importantly, it requires a leadership culture that recognizes that behind every incident ticket is a human being. The future of operations lies in the successful marriage of high-tech intelligence and high-touch empathy. 
  
Future entries will analyze the practical application of AI as a first responder and explore how to monitor team performance without infringing on trust. 

Contact us to modernize payments

Author

Robin Mongeau

Senior Operations Delivery Lead

IBM