Results comparison of HiperSockets versus OSA

Having configured both the OSA configuration and the HiperSockets™ configuration based on the details documented earlier, runs with the set of 45 workload tests were conducted on both configurations. The comparison between the two configurations are shown in the following graphs.

The comparison graphs show the relative differences (in %) between the HiperSockets configuration and the OSA configuration across all simulated user (or uperf connection) counts. The differences for these three key performance metrics are illustrated:

throughput
the raw bandwidth or how much data was transferred in fixed period of time
transaction time
the total or round-trip time required to complete one transaction
CPU efficiency
the amount of CPU resources required to transfer a fixed amount of data

The Y-axis represents the relative difference (in %) for HiperSockets results versus OSA results, while the X-axis lists each user count for a specific workload type. A positive value indicates a stronger HiperSockets performance compared to MacVTap on OSA. A negative value indicates a weaker HiperSockets performance.

The results show that HiperSockets provides better throughput, transaction times and CPU efficiency for 33 of the 45 different workload tests. Two test results were roughly the same, and ten test results were worse than OSA. HiperSockets vs OSA performance will be summarized for each of the five workload types in the following sections.

Highly transactional / small data size tests (rr1c-1x1-users)

Figure 1. Results comparison for very small and medium transactional workload types (rr1c-1x1 and rr1c-200x1000)
Results comparison for very small and medium transactional workload types (rr1c-1x1 and rr1c-200x1000)

The results for this workload type are nearly identical to the highly transactional / medium data size test results. See Figure 1.

Highly transaction / medium data size tests (rr1c-200x1000-users)

See Figure 1.

The following results were observed:

  • At lower levels of concurrency HiperSockets demonstrated extreme throughput advantages (~3.5x better) and significant (70-80%) improvements for transaction times and CPU efficiency.
  • The HiperSockets advantages tend to reduce as concurrency increased.
  • The break-even point for this workload type was at ~50 users/connections.
  • At high levels of concurrency, throughput performance and CPU efficiency can drop as much as ~50% below the levels provided by MacVTap on OSA, with transaction times as much as ~150% worse.

Transactional / large data size tests (rr1c-200x30k-users)

Figure 2. Results comparison for large transactional workload type (rr1c-200x30k)
Results comparison for large transactional workload type (rr1c-200x30k)

The following results were observed:

  • Extreme throughput advantages (more than 4x better) at low concurrency levels with up to a 80% reduction in transaction times and a greater than 2x improvement in CPU efficiency.
  • As seen with the smaller data size transactional workload types, the advantages reduce with increased concurrency.
  • However, even at the highest concurrency levels, HiperSockets continues to demonstrate a 50-125% advantage across the three key metrics over MacVTap on OSA configuration.

Streaming read/large data size tests (str-readx30k-users)

Figure 3. Results comparison for large streaming workload type (str-readx30k)
Results comparison for large streaming workload type (str-readx30k)

The following results were observed:

  • HiperSockets demonstrated advantages across all three key metrics at all tested levels (user counts) of concurrency. Throughput improvements were in the range of 160-325% better.
  • Transaction times were better by 75-90%, and CPU efficiency was 175-300% better.
  • Unlike all other workload types, streaming reads are the only workload type to show steady improvements as the concurrency levels increased.

Streaming write/large data size (str-writex30k-users)

Figure 4. Results comparison for large streaming workload type (str-writex30k)
Results comparison for large streaming workload type (str-writex30k)

The following results were observed:

  • Like streaming reads, streaming write also demonstrated advantages across all metrics at all concurrency levels.
  • Like the transactional workload types, streaming writes throughput and transaction times tended to show a gradual decrease at higher levels on concurrency. The exception to this trend is the 250-user-count result, which showed a marked improvement. This behavior was observed and reproduced in three similar runs.
  • Unlike the streaming reads, streaming write advantages were much more consistent across the different levels of concurrency. Streaming writes achieved a throughput advantage in the range of 70-95%. Transaction times were ~40-47% better, and CPU efficiency was improved by 58-95%.