Sunday, October 4, 2026

Data Guard considerations in RAC environment

The flow of Redo logs Transfer/Apply and commit phase in Data Guard for RAC

Investigating the reasons for Data Guard slowness in an Oracle RAC environment

What are reasons for Data Guard being slow in RAC env?

This statement is important in RAC + Data Guard, because the problem is not simply “Data Guard is slow.” The synchronous redo path can become part of the RAC commit path.

Oracle's documentation explains that with synchronous transport, LGWR can wait for the remote standby write before completing the commit. In RAC, LGWR also has to coordinate with the other RAC instances.

What is actually happening?

A simplified RAC + Data Guard commit path looks like:


 

So, if the network, standby storage, standby server, or redo transport configuration is slow, the effect can eventually appear as higher commit latency on the primary.

Oracle specifically describes this sequence: foreground waits on log file sync, LGWR performs local redo processing, sends redo to a synchronous standby, waits for the remote write, and in RAC also waits for the RAC broadcast acknowledgment.

Redo transport consists of the primary database instance background process sending redo to the standby database background process. You can evaluate whether the network is properly optimized for Oracle Data Guard redo transport. For this, a DBA can use Oracle's oratcptest utility to measure single-stream and multi-stream throughput.

https://docs.oracle.com/en/database/oracle/oracle-database/19/haovw/plan-oracle-data-guard-deployment.html#GUID-CEC2DC1E-4134-4F9C-B0F0-3D9F0D7C5D46

The Oracle utility oratcptest is a general-purpose tool for measuring network bandwidth and latency similar to iperf/qperf which can be run by any OS user.

The oratcptest utility provides options for controlling the network load such as:

  • Network message size
  • Delay time between messages
  • Parallel streams
  • Whether or not the oratcptest server should write messages on disk.
  • Simulating Data Guard SYNC transport by waiting for acknowledgment (ACK) of a packet or ASYNC transport by not waiting for the ACK.

Note: This tool, like any Oracle network streaming transport, can simulate efficient network packet transfers from the source host to target host similar to Data Guard transport. Throughput can saturate the available network bandwidth between source and target servers. Therefore, Oracle recommends that short duration tests are performed and that consideration is given for any other critical applications sharing the same network.

# java -jar oratcptest.jar -help

With asynchronous transport, the primary does not wait for the standby acknowledgment before continuing the local redo processing and, redo data is streamed to the standby in large packets asynchronously.

To tune asynchronous redo transport over the network, you need to optimize a single process network throughput.

If synchronous redo transport is configured, each redo write must be acknowledged by the primary and standby databases before proceeding to the next redo write. You can optimize standby synchronous transport by using the FASTSYNC attribute as part of the LOG_ARCHIVE_DEST setting, but higher network latency (for example, more than 5 milliseconds) impacts overall redo transport throughput. With synchronous transport, the primary must wait for the required standby acknowledgment before the corresponding commit can complete.

 

Understand Your Network Topology for Data Guard

Before troubleshooting Data Guard transport lag, understand the end-to-end infrastructure. Many performance issues attributed to Data Guard are actually caused by the underlying infrastructure.

Document at least:

🔹 Database infrastructure: RAC node count, CPU, memory, and storage I/O
🔹 Network topology: switches, firewalls, routing, and connectivity
🔹 Network capacity: bandwidth and RTT between Primary and Standby
🔹 Redo transport: peak redo generation and per-RAC-instance throughput

Two network-intensive phases deserve special attention:

-           Standby Instantiation
Optimize the degree of parallelism to maximize the throughput when copying database files.

-           Steady-State Redo Transport
Each RAC instance sends redo through its own transport stream. Therefore, single-stream network throughput is critical, not just the aggregate bandwidth.

For symmetric Primary/Standby infrastructure with a properly tuned network capable of handling peak redo generation, Oracle notes that transport lag should generally be less than 1 second.

Tip: Don't troubleshoot Data Guard in isolation. Map the complete path:

Primary → Network → Firewall/Switches → Network → Standby

Then validate that every component can support the required redo transport rate.

 

Assessing and Optimizing Network Performance in Oracle Data Guard

Oracle Data Guard transport performance is not only about bandwidth. The network must be able to sustain the peak redo generation rate of each primary RAC instance.

 

A few key points I consider when assessing a Data Guard network:

🔹 Measure peak redo generation : average AWR rates can hide short workload spikes.
🔹 Test single-stream throughput : each RAC instance ships redo through a single network stream, so per-process throughput matters.
🔹 Check latency and reliability : temporary packet loss, retransmissions, firewalls, or overloaded network devices can quickly create transport lag.
🔹 Validate socket buffers and MTU : OS socket tuning and Jumbo Frames can significantly improve throughput in some environments.
🔹 Be careful with encryption and compression : Oracle Net encryption and COMPRESSION=ENABLE can introduce CPU/throughput overhead. Compression should generally be considered when network bandwidth is the limiting factor.
🔹 Use oratcp for controlled network testing: compare throughput under different MTU, socket-buffer, encryption, and network conditions.

One important sizing principle:

Required network capacity should be based on peak redo generation, with additional headroom, not simply the average redo rate.

For example, if peak redo generation reaches 52 MB/s, designing for approximately 68 MB/s (+30%) provides additional capacity for workload fluctuations.

A healthy Data Guard network should be validated end-to-end: Primary → Network → Standby, rather than assuming that a high-speed network link automatically means sufficient redo transport performance.

 

1. LOG FILE SYNC does not automatically mean "Data Guard problem"

This is an important distinction.

log file sync is the time the foreground waits for LGWR to complete the commit operation.

It can be affected by:

  • local redo I/O
  • CPU scheduling
  • excessive commit frequency
  • RAC inter-instance coordination
  • Data Guard synchronous transport
  • network latency
  • standby redo log I/O

So don't look at:

log file sync = 15 ms, and immediately conclude: Data Guard is causing the problem!

Oracle explicitly warns that log file sync averages can be misleading when evaluating synchronous Data Guard impact.

You need to correlate it with the redo transport waits.

https://docs.oracle.com/en/database/oracle/oracle-database/26/haovw/redo-transport-troubleshooting-and-tuning.html

 

2. The important Data Guard waits

For synchronous transport, I would look at:

·         LNS wait on SENDREQ

·         LNS wait on ATTACH

·         LNS wait on DETACH

Oracle documents these as redo transport wait events. LNS wait on SENDREQ is particularly interesting because it represents time spent waiting for redo data to be written to redo transport destinations.

https://docs.oracle.com/en/database/oracle/oracle-database/26/sbydb/oracle-data-guard-redo-transport-services.html

Also examine:

·         SYNC remote write

and, depending on the version/workload:

·         redo transport related waits

·         log file parallel write

·         log file sync

The key is to determine where the latency is introduced.

 

3. SYNC/AFFIRM can directly increase commit latency

Suppose you have:

LOG_ARCHIVE_DEST_2='SERVICE=STBY SYNC AFFIRM VALID_FOR= (ONLINE_LOGFILES, PRIMARY_ROLE) DB_UNIQUE_NAME=STBY'




The primary waits for the standby to acknowledge that redo has been received and written to persistent standby redo storage.

Therefore:

 

Oracle explicitly states that SYNC/AFFIRM provides stronger protection but introduces performance impact because of the standby redo-log I/O.

This is why standby storage latency matters to primary OLTP response time.

https://docs.oracle.com/en/database/oracle/oracle-database/26/sbydb/data-guard-concepts-and-administration.pdf

4. SYNC/NOAFFIRM is different

In Maximum Availability, you can use:

SYNC/NOAFFIRM, also known as Fast Sync.

The standby acknowledges after receiving the redo rather than waiting for the standby redo log write.

That removes the standby storage write latency from the synchronous acknowledgment path.



This can reduce primary commit latency, but there is a trade-off: in a very specific simultaneous-failure scenario, NOAFFIRM can expose some data that has been acknowledged but not yet persisted at the standby.

Ø  So don't blindly change AFFIRM to NOAFFIRM. It is a protection-vs-performance decision.

 

5. NET_TIMEOUT is extremely important

For synchronous transport, I would always review:

NET_TIMEOUT

Oracle recommends specifying NET_TIMEOUT for synchronous transport because it controls how long LGWR waits for acknowledgment before terminating the redo transport connection.

https://docs.oracle.com/en/database/oracle/oracle-database/21/sbydb/oracle-data-guard-redo-transport-services.html

 

Conceptually:



 

 

 

 

 

 

 


If this is badly configured, a network problem can cause a surprisingly long primary-side stall.

You should therefore review:

SQL> show parameter log_archive_dest;

and specifically inspect:

 

·         SYNC

·         AFFIRM / NOAFFIRM

·         NET_TIMEOUT

·         REOPEN

·         VALID_FOR

·         DB_UNIQUE_NAME

 

6. RAC makes the situation more interesting

This is the part behind Oracle's phrase:

"Avoid cluster related waits"

 

 



Imagine:

 

 

 

 

 

 



A commit can involve both:

·         Data Guard remote synchronization

and:

·         RAC inter-instance synchronization

Oracle's documentation describes LGWR waiting for the synchronous remote write and, for RAC, also waiting for the broadcast acknowledgment from the other instances.

Therefore, you can have a situation where:

Data Guard is healthy + RAC interconnect is slow = higher commit latency

or:

RAC interconnect is healthy + Standby network/storage is slow = higher commit latency or both.

 

7. Network latency is more important than just bandwidth

This is another common mistake.

People often check:

·         10 GbE

·         20 GbE

·         25 GbE

and conclude the network is fast enough.

For synchronous Data Guard, latency is critical.

Oracle's current MAA guidance recommends looking at:

  • network topology
  • network bandwidth
  • network latency
  • firewalls
  • encryption
  • socket buffers
  • MTU
  • peak redo generation rate

and states that a well-tuned network with sufficient bandwidth should normally keep transport lag very low.

https://docs.oracle.com/en/database/oracle/oracle-database/26/haovw/high-availability-overview-and-best-practices.pdf

For example:

Redo generation = 300 MB/s

Network = 10 Gb/s

looks excellent from a bandwidth perspective.

But if:  RTT = 15 ms

and the workload is extremely commit-intensive, the application can still experience significant synchronous commit latency!

 

8. Standby storage can become part of the OLTP critical path

This is probably the most important practical point.

With: SYNC + AFFIRM

your primary transaction response time is partly dependent on:

Standby -> RFS -> Standby Redo Log -> Storage -> ACK

So, if standby storage suddenly changes from:

0.5 ms to: 8 ms

you may see the effect on the primary.

That's why I would monitor standby redo log write latency, not just:

Apply Lag

Transport Lag

 

9. Transport Lag and Apply Lag are not the same

This is another common diagnostic mistake.

Transport Lag = Redo generated on primary

Transport lag represents the amount of redo generated on the primary that has not yet been received by the standby.

but not yet received by standby

while:

Apply Lag   = Redo received

but not yet applied

Oracle explicitly distinguishes these two.

Transport lag is more directly related to the ability to ship redo and therefore RPO; apply lag can additionally indicate an RTO concern.

https://docs.oracle.com/en/database/oracle/oracle-database/23/haovw/high-availability-overview-and-best-practices.pdf

 

You can therefore have:

Transport Lag = 0, Apply Lag = 30 sec

meaning:

Network/transport is keeping up, but standby apply is behind.

That's different from:

Transport Lag = 30 sec, Apply Lag = 30 sec

which points much more strongly toward a transport bottleneck.

 

10. What I would check in a RAC/Data Guard performance investigation

My checklist would be:

Primary RAC

    §    log file sync

    §  log file parallel write

    §  redo write time

    §  redo generation rate

    §  commit rate

    §  DB CPU

    §  CPU scheduling

    §  RAC interconnect latency

    §  GC-related waits

 

Data Guard transport

    §  LNS wait on SENDREQ

    §  LNS wait on ATTACH

    §  SYNC remote write

    §  transport lag

    §  redo transport throughput

    §  redo generation vs transport throughput

Network

    §  RTT latency

    §  packet loss

    §  bandwidth

    §  socket buffers

    §  MTU

    §  firewall

    §  network encryption

    §  network congestion

Standby

    §  RFS activity

    §  standby redo log I/O latency

    §  standby redo log configuration

    §  storage latency

    §  CPU

    §  I/O saturation

    §  Redo Apply rate

    §  Apply Lag

    §  Transport Lag

Configuration

    §  SYNC / ASYNC

    §  AFFIRM / NOAFFIRM

    §  NET_TIMEOUT

    §  REOPEN

    §  DATA_GUARD_SYNC_LATENCY

    §  standby redo logs

    §  real-time apply

    §  Data Guard Broker configuration

Oracle also notes that Data Guard automatically tunes redo transport, but network, storage, FRA, and redo-transport configuration can still be tuned when required.

https://docs.oracle.com/en/database/oracle/oracle-database/26/sbydb/oracle-data-guard-redo-transport-services.html

The key RAC + Data Guard relationship

I would summarize the original Oracle statement like this Figure:

 


Every additional millisecond introduced into the synchronous commit path has the potential to affect OLTP commit response time.

And that's why simply saying:

"Data Guard is synchronized and Transport Lag is 0"

does not prove that Data Guard is optimally tuned.

You need to prove that the redo transport path is not becoming the bottleneck in the RAC commit path.

  

Appendix

I’d make it more lab-oriented: show how to measure peak redo, calculate the required throughput, inspect the network, then validate socket buffers and MTU with Oracle’s oratcptest. Current guidance also uses 3× BDP (BDP = Bandwidth × RTT) as a practical upper target for TCP socket buffers on high-latency/high-bandwidth paths.

BDP tells you how much data needs to be "in flight" on the network to fully utilize a link:

For example:

  • Network = 1 Gbit/s and RTT = 20 ms

Convert bandwidth: 1 Gbit/s ÷ 8 = 125 MB/s

Then:       BDP = 125 MB/s × 0.020 s = 2.5 MB

So, the network can have approximately 2.5 MB of data in flight during one RTT.

What does 3× BDP mean?

Simply: 3 × BDP = 3 × 2.5 MB = 7.5 MB

So, you might test TCP socket buffers around 7.5 MB or higher.

The reason for using a multiple such as 3× is to provide enough TCP buffering to keep the link busy despite network timing variations, congestion, retransmissions, and TCP behavior.

 

Why is this important for Data Guard?

Imagine your Data Guard network is:

Primary RAC --à-------1 Gbit/s ----- RTT = 20 ms -------------à Standby

 

If the TCP buffers are too small, the sender may not be able to keep enough redo data in flight before waiting for acknowledgements.

You could have:

Network capacity:       125 MB/s

Peak redo generation:    80 MB/s

but still experience poor redo transport throughput because of insufficient TCP buffering.

That's particularly relevant for high-bandwidth / high-latency Data Guard links.

Important distinction

3× BDP is not "set the network bandwidth to 3×."

It means approximately:

 

TCP socket buffer ≈ 3 × (Bandwidth × RTT)

 

For example:

Bandwidth

RTT

BDP

3× BDP

1 Gbit/s

1 ms

125 KB

375 KB

1 Gbit/s

10 ms

1.25 MB

3.75 MB

1 Gbit/s

20 ms

2.5 MB

7.5 MB

10 Gbit/s

20 ms

25 MB

75 MB

10 Gbit/s

50 ms

62.5 MB

187.5 MB

 

This is why RTT becomes extremely important for Data Guard across geographically separated sites.

Oracle Data Guard: Measuring and Optimizing Redo Transport Network Performance

When investigating Data Guard transport lag, one of the first questions I ask is:

Can the network transport redo faster than the primary generates it?

A simple bandwidth test such as iperf3 is useful, but Data Guard has an important characteristic: each primary RAC instance ships redo through its own transport stream. Therefore, single-stream throughput matters, not only aggregate network bandwidth.

1- Measure the actual peak redo rate

Instead of relying only on 30/60-minute AWR averages, I prefer measuring redo generation over individual archived logs:

SELECT

    THREAD#,

    SEQUENCE#,

    BLOCKS * BLOCK_SIZE / 1024 / 1024 AS REDO_MB,

    (NEXT_TIME - FIRST_TIME) * 86400 AS SECONDS,

    (BLOCKS * BLOCK_SIZE / 1024 / 1024) /

    ((NEXT_TIME - FIRST_TIME) * 86400) AS REDO_MB_SEC

FROM V$ARCHIVED_LOG

WHERE (NEXT_TIME - FIRST_TIME) * 86400 <> 0

  AND FIRST_TIME BETWEEN

      TO_DATE('2026/10/03 08:00:00','YYYY/MM/DD HH24:MI:SS')

  AND TO_DATE('2026/10/03 12:00:00','YYYY/MM/DD HH24:MI:SS')

  AND DEST_ID = 1 ORDER BY FIRST_TIME;

 

 

For example, if the peak rate is 52 MB/s per RAC instance, I would not design the transport network for exactly 52 MB/s.

With 30% headroom:

52 × 1.3 ≈ 68 MB/s

For multiple RAC instances, evaluate the requirement per instance and for the aggregate path.

2- Check average redo write size

SELECT

    NAME,

    VALUE

FROM V$SYSSTAT

WHERE NAME IN ('redo size', 'redo writes');

Then:

Average Redo Write Size =

    REDO SIZE / REDO WRITES

This number is useful when evaluating MTU because the optimal network behavior depends on the actual redo message characteristics.

3- Calculate the TCP Bandwidth-Delay Product

For example:

Network bandwidth = 1 Gbit/s

RTT = 20 ms

BDP = 1,000,000,000 / 8 × 0.020 = 2.5 MB

Oracle recommends considering socket-buffer sizes based on BDP; current Oracle guidance indicates 3× BDP as a practical target for high-latency/high-bandwidth environments.

Check Linux:

sysctl net.ipv4.tcp_rmem

sysctl net.ipv4.tcp_wmem

sysctl net.core.rmem_max

sysctl net.core.wmem_max

Example test values:

sysctl -w net.ipv4.tcp_rmem='4096 87380 16777216'

sysctl -w net.ipv4.tcp_wmem='4096 16384 16777216'

Do not blindly copy these values into production—the correct values should be determined through testing against the actual bandwidth and RTT. Oracle also recommends making tested values persistent in /etc/sysctl.conf.

4- Test the network with Oracle oratcptest

Rather than testing only with generic network tools, use Oracle's oratcptest to evaluate TCP throughput and socket-buffer behavior.

On the standby:

java -jar oratcptest.jar -server [IP of standby host or VIP in RAC configurations] -port=<any available port number>

Then from the primary, connect to the standby and test the transport path.

Run the test client. (Change the server address and port number to match that of your server started)

$ java -jar oratcptest.jar [IP of standby host or VIP in RAC configurations]  -port=<port number> -mode=async -duration=120 -interval=20s

This process can be scheduled to run at a given frequency using the -freq option to determine if the bandwidth varies at different times of the day. For instance setting -freq=1h/24h will repeat the test every hour for 24 hours.

Run the test with different socket-buffer sizes and compare:

Socket Buffer     Throughput

-------------     ----------

1 MB              ...

4 MB              ...

8 MB              ...

16 MB             ...

32 MB             ...

The objective is to find the point were increasing the socket buffer no longer produces meaningful throughput improvement.

Oracle notes that oratcptest reports approximately half of the socket buffer allocated to the socket, so interpret the reported value accordingly.

5- Test MTU 1500 vs 9000

First establish the current MTU:

ip link show

or:

ip addr show

Then test the path:

ping -M do -s 8972 <standby-ip>

If the network is designed for Jumbo Frames, test MTU 9000 end-to-end.

For example:

ip link set dev bond0 mtu 9000

Then repeat the oratcptest measurements.

The important point is measurement, not assumption: MTU 9000 is not automatically faster. Oracle specifically recommends comparing the throughput with the existing MTU and a larger MTU such as 9000.

6- Check Data Guard configuration

Finally, check whether compression or Oracle Net encryption is affecting the transport path:

SELECT DEST_ID,

       DEST_NAME,

       STATUS,

       TARGET,

       TRANSMIT_MODE,

       COMPRESSION,

       NET_TIMEOUT

FROM V$ARCHIVE_DEST

WHERE TARGET = 'STANDBY';

For example:

LOG_ARCHIVE_DEST_2 =

'SERVICE=STBY

 ASYNC

 NOAFFIRM

 COMPRESSION=DISABLE

 

Compression can help when bandwidth is the bottleneck, particularly on low-bandwidth/high-latency links, but it also introduces CPU processing. Oracle therefore recommends evaluating it rather than enabling it automatically.

 

The practical workflow

My Data Guard network assessment is therefore:

Peak Redo Rate → RTT → BDP → Socket Buffers → Single-Stream Throughput → MTU → Encryption/Compression → Transport Lag

The final question is simple: Can each primary instance continuously ship its peak redo generation rate, with sufficient headroom, across the real Data Guard network path?

If the answer is no, eventually the network becomes the bottleneck, and transport lag is only the symptom.

 Ref:

https://docs.oracle.com/en/database/oracle/oracle-database/21/haovw/configure-and-deploy-oracle-data-guard.html

https://docs.oracle.com/en/database/oracle/oracle-database/21/sbydb/oracle-data-guard-redo-transport-services.html

https://docs.oracle.com/en/database/oracle/oracle-database/21/haovw/plan-oracle-data-guard-deployment.html

https://docs.oracle.com/en/database/oracle/oracle-database/21/haovw/configure-and-deploy-oracle-data-guard.html

 

 

 

 


No comments:

Post a Comment

Data Guard considerations in RAC environment

The flow of Redo logs Transfer/Apply and commit phase in Data Guard for RAC Investigating the reasons for Data Guard slowness in an Oracle...