Skip to main content

Winning the AI race in refining

 

Published by
Hydrocarbon Engineering,

In this special report, Aashna Punwani, UptimeAI, discusses how AI tools can be used to speed up decision-making within refinery and petrochemical operations.

Over the last few years, it feels as though the entire world has entered the ‘AI race’, with leaders across industries being pushed to make their organisations more AI-enabled and identify where AI can create lasting business impact. The oil and gas industry is no exception. However, for downstream operators, the decision of where and how to apply AI is unique. Unlike many industries that can realise value through productivity improvements alone, the real financial levers in refining operations depend on optimising complex physical processes and extracting more value from existing assets.

This raises the main question downstream oil and gas operators are asking themselves: how can AI reduce margin leakage? Margin leakage, the gap between a refinery’s theoretical economic potential and realised performance, represents one of the largest economic opportunities in downstream operations. Even recovering a small percentage of lost margin can translate into millions of dollars of annual value for a refinery. There are many contributors to margin leakage to explore, including energy inefficiency, process optimisation and control, and unplanned downtime due to reliability issues. The unlock comes from exploring them together, rather than in isolation.

Margin leakage is not solved in silos

Traditional asset performance management (APM) and predictive analytics solutions aim to maximise equipment reliability without business and supply chain context. Real-time optimisation (RTO) and advanced process control (APC) solutions aim to maximise process economics without knowledge of changing reliability issues or process constraints. Connecting the dots in this world of competing objectives has historically been left to a pool of experienced operators and engineers. But this pool is growing smaller every day.

The alerts, recommendations, and insights these systems generate ultimately rely on human experts to interpret them, decide on a course of action, and execute it. While insights can be valuable, information alone does not create value. Value is realised when decisions are made and actions are implemented. The gap between when an insight is surfaced and a decision is made – where an expert is injected into the process to investigate, reason, apply experience, and issue a course of action – is called decision latency.

In downstream operations, this decision latency can significantly reduce the economic impact of AI. Many opportunities are time-sensitive and lose value as operating conditions change. For example, an optimisation model may recommend adjusting a reactor to a more economically optimal temperature, but by the time the recommendation is reviewed, approved, and implemented, the crude slate or process conditions may have changed and the recommendation is no longer optimal. Similarly, a predictive maintenance alert may identify elevated bearing vibration on a critical compressor. However, if several days pass while an expert reviews the alert, analyses historical trends, creates a work request, and schedules inspection, damage may already have accumulated, ultimately reducing equipment life and increasing future maintenance costs.

These examples highlight a key limitation of advisory systems: the dependency on human decision cycles can cap the value that AI ultimately delivers. The next evolution of industrial AI is not only generating better insights; it is reducing the time between detection, decision, and action.

Closing the insight to execution gap

More accurate detection models and generative AI retrieval have helped provide insights with improved context but have failed to reach the next rung on the decision ladder. The next major advancement in industrial AI is not a better prediction model, and it is not a large language model (LLM). It is the ability to reason in the same way that experienced plant experts think through a situation, analysing the full system context, and then make a decision on the optimal course of action.

A new class of industrial AI known as AI reasoning agents have emerged to narrow the gap between insight and action. Like their predecessors, these systems start with an insight. But unlike technologies designed with insight generation as the end goal, reasoning agents go the next step to decision and execution. Reasoning agents operate like domain specialists embedded directly into operational workflows. They take on that human interpretation responsibility, eliminating the decision latency by taking in available data and knowledge and applying the correct combination of skills – for example sensor trend analysis, piping and instrumentation diagram (P&ID) interpretation, failure mode and effects analysis (FMEA), and criticality analysis – reasoning to determine the best path forward (Figure 1). These agents can operate in an advisory capacity, providing a less experienced engineer or operator with a real-time, expert-level decision, and even execute some changes within guardrails.


Figure 1. Reasoning agents focused on objectives orchestrate the correct combination of skills, context, and data for optimal, real-time decision making.

AI reasoning agent orchestration

Reasoning agents are built on a data foundation that ingests relevant contextual information that has historically existed in siloed systems. This can include large volumes of operational data across historians, control systems, maintenance records, operating procedures, engineering documentation, and economic constraints. The knowledge graph foundation provides accurate context to the interconnected data sources.

The agents come pre-built with the relevant domain skills to solve specific problems. A root cause agent might analyse sensor trends, maintenance and criticality, manuals, standard operating procedures (SOPs), documents, and P&IDs as well as map failure modes and compare amongst similar assets. A maintenance optimisation agent might analyse asset criticality, maintenance history, manuals, SOPs, past root cause analyses (RCAs), and run what-if scenarios. It takes in the relevant information, applies domain-specific skillsets, and synthesises the data to make an informed, optimal decision, just like human experts would.

For example, if a compressor vibration signal begins increasing, a traditional predictive maintenance system may simply issue an alert and rely on a reliability engineer to investigate the source of the vibration. A reasoning agent detects the abnormal vibration, correlates with process changes, evaluates whether operations changed, reviews maintenance history, compares behaviour against historical events, estimates risk progression, quantifies economic exposure, then recommends inspections or the appropriate path forward. Similarly, in process optimisation, rather than recommending a one-time reactor temperature adjustment, a reasoning agent can continuously monitor changing feed conditions, operating constraints, and economics to determine whether the recommendation remains optimal, then adapt to the new optimum in real time.

Case studies demonstrate elevated potential of reasoning agents

Reasoning agents have proven an effective lever for reducing margin leakage linked to process, reliability, and other issues. Leveraging cross-functional domain skillsets with real-time and historical data, they are able to make timely decisions that move the needle on business metrics. The following case studies demonstrate the value potential in the downstream oil and gas industry.

Case study 1

Challenge

At a large refinery, operations teams were experiencing repeated feed pump reliability issues that disrupted production and periodically required shutdowns for clean-up and maintenance. Despite regular monitoring and investigation, the root cause remained difficult to isolate because the failure mechanism extended beyond the pump itself and involved interactions across multiple process units.

At the time, traditional monitoring tools showed the pump operating within expected limits and alarm thresholds had not been exceeded. However, process conditions upstream had begun to deteriorate. Fouling in a reboiler had reduced heat transfer performance, preventing heavier crude fractions from being adequately heated before entering the distillation process. This created localised temperature effects that promoted polymerisation. As operating conditions evolved, the feed pump was required to work harder to maintain throughput, increasing mechanical load and creating conditions that would eventually lead to equipment failure.

Solution

To address the challenge, the refinery deployed a root cause agent designed to reason across process and reliability data and alerts, and get to the source of issues rather than the symptoms.

The agent continuously analysed relationships across the distillation column, reboiler, and feed pump, constructing a causal chain on how the issue progressed across the system. It first identified abnormal behaviour consistent with declining reboiler performance, then connected subsequent shifts in temperature conditions with increased likelihood of polymerisation, and finally linked those process changes to rising mechanical stress on the downstream feed pump.

By presenting engineering teams with an assembled investigation hypothesis (Figure 2) and recommended actions rather than just raw signals, the system reduced the time required to move from anomaly detection to operational action.


Figure 2. Root cause agent issues a hypothesis with transparent causal chain and actionable recommendations.

Results

The earlier diagnosis gave operations teams time to intervene before the issue progressed into a mechanical failure event. Maintenance activities were planned around restoring reboiler performance, preventing continued polymerisation, and minimising the risk on the feed pump.

This individual intervention avoided a >US$1 million equipment failure and prevented the operational disruption associated with an unplanned shutdown. More importantly, the event demonstrated that the limiting factor was not the availability of data, it was the speed at which disparate information could be converted into an actionable understanding – one that treated the problem at its true root cause.

Case study 2

Challenge

At a downstream oil and gas facility, the preventive maintenance (PM) programme for rotating equipment had become stagnant. It was standardised at the equipment class level, meaning pumps received identical inspection frequencies, oil checks, and component replacement schedules regardless of differences in service, utilisation, operating severity, process conditions, or failure history.

One critical charge pump emerged as a representative example. The asset had accumulated a mix of preventive and corrective maintenance (CM) work over time, but the PM strategy remained largely unchanged. Reliability teams suspected opportunities existed to improve maintenance effectiveness, but evaluating a single asset required manually reviewing work orders, maintenance records, failure history, and operating data across multiple systems. Then they would have to repeat the process for every asset they wanted to optimise maintenance for. The effort required meant these optimisation exercises rarely occurred.

Solution

To address this challenge, the facility deployed a maintenance optimisation agent that continuously evaluated maintenance effectiveness at the individual equipment level, not the asset class level.

The agent mirrored the existing maintenance hierarchy and reviewed the full preventive and CM history of the selected pump alongside operating conditions and available sensor data. Using reliability analysis methods to evaluate whether observed failure patterns reflected wear-out behaviour, random failures, or over-maintenance, the agent generated a prioritised set of recommendations.

For the selected pump, recommendations included extending maintenance intervals where historical performance showed low corrective risk, increasing maintenance frequency where corrective events exceeded expected levels, introducing new preventive activities where recurring issues lacked PM coverage, and removing calendar-based activities where condition monitoring already provided sufficient visibility (Figure 3).


Figure 3. Maintenance optimisation agent issues optimal maintenance interval decisions balancing PM and CM costs.

Approved recommendations were structured for direct integration into existing maintenance workflows rather than requiring manual redesign of maintenance plans.

Results

For the initial pump scope, the revised maintenance strategy identified approximately US$100 000 in annual maintenance savings through a combination of reduced unnecessary preventive work and lower CM exposure.

Following the initial results, the optimisation approach was expanded to four other rotating equipment assets across the site, and the identified savings grew to approximately US$300 000.

The broader lesson was not that maintenance budgets should be reduced. Rather, the case demonstrated that maintenance strategies themselves can become sources of margin leakage when they remain static while operating conditions evolve. By continuously reassessing maintenance decisions at the asset level, reasoning agents can reduce the latency between changing asset behaviour and maintenance action.

AI that reasons will win the AI race

This shift from insights to decisions fundamentally changes how value is created by technology for downstream operations. Traditional AI systems improve information availability but remain constrained by human throughput and decision latency. Reasoning agents increase capacity by compressing the time required to move from observation to action. Engineers and operators remain responsible for oversight, but instead of spending time collecting data and investigating, they engage with higher-value decisions faster. Insights spark investigations, but decisions move the needle on margins.

 

This article has been tagged under the following:

Downstream news