Skip to main content

When Vendor Defaults Hide the Real Capability of a System

Based on a LinkedIn post originally published on 27 March 2026

While working in New Zealand, my employer received several IBM enterprise-storage systems.

The platforms were shared across multiple technology groups:

  • the mid-range UNIX team;
  • the Windows environment;
  • and the mainframe organisation.

A shared storage platform serving such different consumers needs to satisfy a wide range of performance, consistency, resilience and compatibility requirements.

I wanted to understand how the systems behaved in practice, so I began running performance tests and comparing the results with other storage platforms available in the environment.

The results were puzzling.

Peak performance appeared to be capped at approximately the I/O throughput of a single Fibre Channel disk.

For a storage system of that class, this made little sense. The hardware contained multiple drives, controllers, cache and paths. The measured result did not reflect the parallel capability suggested by the physical platform.

When the architecture and measurement disagree

An enterprise-storage system is designed to combine the capacity and performance of many physical devices.

At a simplified level, the expected path looked like this:

Host workload
      |
Multiple Fibre Channel paths
      |
Storage controllers and cache
      |
Many physical disks
      |
Expected parallel I/O capability

The measured behaviour looked more like this:

Host workload
      |
Enterprise storage system
      |
Apparent throughput ceiling
      |
Approximately one disk's I/O rate

When the architecture implies parallelism but the measurement behaves serially, something is constraining the path.

The constraint might be technical:

  • one saturated host path;
  • an incorrectly configured multipath policy;
  • a controller bottleneck;
  • a cache policy;
  • a narrow storage layout;
  • or a workload that cannot generate sufficient parallelism.

It might also be intentional.

The answers did not explain the behaviour

I continued asking IBM for an explanation.

The initial responses were vague, incomplete or redirected toward other parts of the environment. Nobody provided a clear account of why a platform with substantial internal resources consistently produced such a specific ceiling.

This situation is common in complex infrastructure. A support interaction may concentrate on whether the system is functioning according to specification, while the customer is asking why the specification produces an unexpected operational result.

Support question:
Is the system operating as configured?

Engineering question:
Why was it configured to behave this way?

Business question:
Does that behaviour serve our workload?

All three questions are legitimate, but they require different evidence.

The engineer who finally explained it

Eventually, an IBM engineer visited our office.

I used the opportunity to ask questions—persistently, but respectfully. After enough examination, the explanation emerged.

The performance limit was not a hardware fault.

It was a design choice.

Each logical unit was constructed using disks from one disk shelf rather than distributing the LUN across drives located in several shelves.

STANDARD LUN LAYOUT

Disk shelf A
+----+----+----+----+
| D1 | D2 | D3 | D4 | ----> LUN A
+----+----+----+----+

Disk shelf B
+----+----+----+----+
| D5 | D6 | D7 | D8 | ----> Other LUNs
+----+----+----+----+

One LUN remains bounded by the resources
available within one shelf.

A design intended to increase parallelism might instead distribute a logical unit across physical resources associated with several shelves:

PARALLEL LAYOUT CONCEPT

Disk shelf A                 Disk shelf B
+----+----+                   +----+----+
| D1 | D2 |                   | D5 | D6 |
+----+----+                   +----+----+
      \                          /
       \                        /
        +------ One LUN -------+
          Broader parallelism

The exact suitability of either arrangement depends on the storage architecture, protection model and workload. The important discovery was that the observed ceiling followed the chosen layout rather than the apparent capability of the complete hardware platform.

Consistency was the product decision

When I asked why the LUNs were not distributed more broadly, the answer was unexpected:

“We use a standard approach to avoid different customer experiences.”

Performance consistency had been prioritised over maximum performance potential.

From a vendor perspective, this has understandable advantages:

  • deployments behave more consistently;
  • support teams encounter fewer customer-specific layouts;
  • capacity planning follows a familiar model;
  • documentation is easier to standardise;
  • and one customer does not receive dramatically different behaviour merely because an implementation engineer selected another layout.

The default reduced variability.

It also concealed performance that the underlying platform might have delivered under a different design.

A vendor default is often an optimisation—but the variable being optimised may be supportability, consistency or risk rather than your workload’s performance.

The cache behaviour revealed another assumption

I then asked why the storage cache was not absorbing or smoothing the heavy I/O load more effectively.

The explanation exposed another architectural priority.

The cache was synchronised with disk by design to preserve strict write behaviour required by mainframe workloads and demanding systems operating against the shared platform.

Again, this was not evidence that the storage system was defective. It reflected the guarantees the platform had been configured to provide.

Performance-oriented expectation:

Host write
    |
Storage cache accepts write
    |
Host continues
    |
Disk synchronisation follows


Stricter durability behaviour:

Host write
    |
Storage cache and disk state coordinated
    |
Required persistence condition satisfied
    |
Completion acknowledged

Stronger write guarantees can introduce latency or limit how aggressively cache absorbs work. For workloads whose correctness depends on those guarantees, the trade-off is entirely reasonable.

For other consumers, the same default may feel like an unexplained performance restriction.

Shared platforms accumulate hidden assumptions

The IBM systems served UNIX, Windows and mainframe consumers.

Those workloads did not necessarily share the same priorities:

WORKLOAD                  POSSIBLE PRIORITY
------------------------  --------------------------------
Mainframe transaction     Strict write semantics
UNIX database             Parallel throughput and latency
Windows file services     Capacity and predictable response
Large Sun platform        Sustained high I/O rates
Shared operations team    Standardisation and supportability

A single standard configuration inevitably embodies a compromise.

The danger is not the compromise itself. Shared platforms require compromise. The danger is when the assumptions remain invisible and every consumer is told that the default represents the platform’s natural capability.

Once the design choice became visible, the measured result was no longer mysterious.

Performance ceilings can be policy ceilings

Engineers often begin a performance investigation by looking for faults:

  • a saturated device;
  • an incorrect parameter;
  • a firmware defect;
  • a broken path;
  • or inefficient application behaviour.

But a stable and repeatable ceiling may indicate an intentional boundary.

The system may have been designed to limit concurrency, preserve fairness, enforce a durability condition or create consistent behaviour across customers.

Observed ceiling
      |
      +---- Hardware limit?
      |
      +---- Configuration limit?
      |
      +---- Workload limit?
      |
      +---- Protection requirement?
      |
      +---- Fairness or policy limit?
      |
      +---- Vendor support model?

Understanding which boundary is active determines whether changing it is possible, safe and worthwhile.

Measure the system you actually received

Product specifications describe potential capability under defined conditions. They do not guarantee that an installed system has been configured to expose every dimension of that capability to every workload.

Validation should therefore begin with measurement.

A useful assessment examines:

  • throughput and latency across several workload types;
  • behaviour as concurrency increases;
  • path utilisation;
  • cache effectiveness;
  • performance during degraded operation;
  • the relationship between logical volumes and physical resources;
  • and the guarantees associated with write completion.

The objective is not to prove that the vendor is wrong. It is to establish how the delivered system behaves under the organisation’s actual workloads.

Question defaults without discarding their purpose

A default should not be changed merely because greater performance appears possi

AI Assistance Disclosure

This article is based on my original ideas, experience, analysis and conclusions. Artificial intelligence tools were subsequently used as editorial and research assistants to review grammar and wording, improve structure and presentation, organise some arguments into clearer logical sections, and help review references to legal, regulatory and technical concepts.

Where relevant, factual and regulatory references were checked against the sources cited in the article. AI assistance does not replace professional legal, regulatory, financial or technical advice, and the final selection, interpretation, opinions and conclusions presented here remain my own.

Comments

Popular posts from this blog

Movies - The Bubble (2022)

  Back to Evolution (2001) .

IT - Fixing Windows Error 1327: Account Restrictions Are Preventing This User from Signing In

Fixing Windows Error 1327: Account Restrictions Are Preventing This User from Signing In Introduction Error 1327, “Account restrictions are preventing this user from signing in,” is a perplexing and disruptive issue that occurs on some Windows 10 and Windows 11 machines. The message typically appears at login or while connecting to remote resources, like shared folders, network drives, or remote desktops. Table of Contents Symptoms of Error 1327 Common Causes Step-by-Step Troubleshooting Advanced Fixes Automation via PowerShell Prevention Tips Further Reading Symptoms of Error 1327 Users experiencing this error may encounter one or more of the following: Login screen fails after credentials are entered. Error message appears when accessing mapped drives or network resources. Remote Desktop Connection (RDP) is rejected with the 1327 message. Group Policy logon restrictions silently block access. Co...

ATA Drive Capacity Limitations

ATA interface versions up through ATA-5 suffered from a drive capacity limitation of about 137GB (billion bytes). Depending on the BIOS used, you can further reduce this limitation to 8.4GB, or even as low as 528MB (million bytes). This is due to limitations in both the BIOS and the ATA interface, which when combined create even further limitations. To understand these limits, you have to look at the BIOS (software) and ATA (hardware) interfaces together. NOTE In addition to the BIOS/ATA limitations discussed in this section, various operating system limitations exist. These are described later in this chapter. The limitations when dealing with ATA drives are those of the ATA interface as well as the BIOS interface used to talk to the drive. A summary of the limitations is shown in Table 7.12. Table 7.12. ATA/IDE Capacity Limitations for Various Sector Addressing Methods Sector Addressing Method Total Sectors Calculation Maximum Total Sectors Maximum Capacity (Byte...