Validation and evidence

A100 and laptop testing now establish two bounded S2 evidence positions.

A self-run Google Colab A100 fixed-demand POC completed three valid matched pairs through vLLM. Separately, the completed RTX 5050 laptop programme established cross-model mechanism evidence, a bounded Llama 3.1 capacity result and endurance. The two evidence strata are reported separately and are not pooled.

Two matched GPU test lanes feeding telemetry and evidence into a shared audit package.
Native and S2 are tested under matched conditions and the same acceptance requirements.
01

How to interpret the evidence

Established inside the test boundary. Not generalised beyond it.

Quality-accepted service means a delivered outcome that passes the agreed quality and service requirements. The completed programme measures that useful outcome separately from the fresh physical execution required to provide it.

The evidence establishes

  • A positive indicative matched-pair effect on a full NVIDIA A100-SXM4-80GB through vLLM 0.23.0 using Qwen2.5-7B-Instruct BF16.
  • Three of three valid A100 fixed-demand pairs passed at approximately 24.99 logical requests per second.
  • Across those A100 pairs, Full S2 maintained effectively equivalent quality-accepted service while using 26.19% fewer physical calls and 1.76% less GPU-board energy per accepted request.
  • Mechanism evidence across four qualified Ollama models.
  • Five valid fixed-demand pairs that all favoured S2 for physical-call reduction and GPU-board energy per accepted request.
  • A bounded capacity result for Llama 3.1 8B Q4_K_M through Ollama 0.31.1 on the tested RTX 5050 laptop.
  • Six balanced one-hour confirmation trials for the bounded Llama 3.1 comparison.
  • Four-hour S2 endurance under the tested conditions.

The evidence does not establish

  • Universal improvement across workloads.
  • Production readiness or production-cluster validation.
  • Independent validation.
  • A measured A100 capacity uplift, native capacity ceiling or maximum sustainable uplift.
  • A causal A100 efficacy claim: the pair-only POC used static controls and did not run measured unique-work negative controls or sensitivity conditions.
  • Whole-system, rack, cooling or facility energy savings.
  • Universal latency improvement.
  • Applicability to an individual customer without customer-specific evidence.

The A100 result is a descriptive, indicative matched-pair observation. It is not causal efficacy or a capacity-ceiling result. No customer-specific conclusion follows without customer-specific evidence.

02

Measurement and assurance

Measure accepted service and the physical work required to deliver it.

Unsupported measurements are not guessed. Each conclusion remains inside the boundary of the evidence captured during a valid comparison.

Quality-accepted logical requests

The number of useful service obligations completed within the agreed quality rules.

Physical model calls

How much fresh model execution was actually required to serve that accepted demand.

Latency and queue stability

Whether service remains responsive and demand remains controlled throughout a valid run.

GPU-board energy

Energy reported at the GPU board boundary for the tested workload.

Hardware consistency

Whether the compared runs used a stable, comparable hardware environment.

Decision and validation outcomes

The evidence needed to explain which work was served, forwarded or excluded.

Operational assurances

What remains true when S2 is present

  • The model obligation remains unchanged.
  • Required work is forwarded.
  • Reused or shared results remain quality-gated.
  • Unsafe or unavailable optimisation fails open.
  • Workload and tenant boundaries are respected.
  • Decisions are retained for evidence.
  • Unsupported measurements are reported as unavailable.

Explicit claim boundaries

What S2 does not claim.

Commercial credibility depends on stating the boundary as clearly as the opportunity.

03

Evidence process

Two completed evidence strata. No cross-boundary pooling.

The A100 POC and laptop programme answer different questions. Both retain quality-accepted service, physical-call evidence, queue behaviour, GPU-board telemetry and hardware state. Only the laptop Llama 3.1 capacity campaign established a directly bounded native ceiling.

Completed A100 fixed-demand POC

Positive indicative paired effect observed.

3 of 3 pairs passed

In one self-run Google Colab session using a full NVIDIA A100-SXM4-80GB, Qwen2.5-7B-Instruct BF16 and vLLM 0.23.0, Full S2 maintained effectively equivalent quality-accepted service while reducing physical model calls and GPU-board energy per accepted request.

26.19% fewer

Physical vLLM calls
Observed across three valid fixed-demand pairs at approximately 24.99 logical requests per second.

1.76% lower

GPU energy per accepted request
Measured at the GPU-board boundary across the same three A100 pairs.

3 of 3

Matched pairs passed
Every pair retained effectively equivalent accepted service with stable queues.

100%

Full S2 quality acceptance
44,976 of 44,976 scheduled Full S2 requests were accepted under the tested requirements.

9.72% lower

P50 latency
The aggregate mean moved from 0.31243 seconds natively to 0.28208 seconds with Full S2.

0

Final queue depth
Every native and Full S2 condition ended with no unfinished queue.

Aggregate mean across three valid matched pairs
MetricNativeFull S2
Scheduled requests44,97644,976
Quality-accepted requests44,97544,976
Physical model calls44,97633,198
GPU-board energy367,284.988 J360,824.630 J
Energy per accepted request8.16643 J8.02260 J
Energy per accepted token0.67661 J0.66476 J
P50 latency0.31243 s0.28208 s
P95 latency0.61651 s0.61088 s
P99 latency0.68062 s0.67845 s
Service-level attainment99.9978%100%
Consistency across the three matched pairs
PairFewer callsLower J/requestFinal-window energy
Pair 126.04%1.75%1.59%
Pair 226.28%1.16%1.44%
Pair 326.24%2.36%4.11%

Pair 1 retained one native summarisation rejection in its physical-work and energy accounting. Full S2 accepted every scheduled request.

Test environment and stability

Hardware
Full NVIDIA A100-SXM4-80GB; MIG disabled
Model
Qwen2.5-7B-Instruct BF16
Serving environment
vLLM 0.23.0 on Google Colab
Fixed demand
Approximately 24.99 logical requests per second
Stability
Zero timeouts or unfinished requests; every condition ended at queue depth zero
Safety and telemetry
No thermal throttling, ECC events, contamination or telemetry failures

Physical-call accounting

11,778 fewer physical calls, reconciled exactly.

11,750Completed-response reuse
28Shared generation
0Obsolete-work removal

These are observed service-accounting categories. They do not describe the protected eligibility or decision logic.

Evidence integrity

Master archive and pair downloads verified.

98,027 internal SHA-256 manifest entries matched, and 3 independently downloaded pair archives matched the master evidence.

Master SHA-256845392a798dae49dc7231f6840e742646209c0a830fb28e46ca4e80e981bebc5

Completed RTX 5050 laptop programme

Cross-model mechanism, bounded capacity and endurance evidence.

Completed

This separate Ollama evidence stratum retains the previously completed five-pair cross-model comparison and directly bounded Llama 3.1 capacity result. It is not pooled with the A100 POC.

28.19% fewer

Physical model calls
Aggregate result across five valid fixed-demand crossover pairs covering four qualified Ollama models.

16.10% lower

GPU energy per accepted request
Measured at the GPU-board boundary across the same five matched pairs.

At least 50.09% more

Quality-accepted logical workload
Directly bounded for Llama 3.1 8B Q4_K_M through Ollama 0.31.1 on the tested RTX 5050 laptop.

5 of 5

Valid pairs favoured S2
Every valid pair showed fewer physical calls and lower GPU energy per accepted request and token.

6 x 1 hour

Balanced confirmation trials
Three native and three S2 trials confirmed the bounded Llama 3.1 capacity result.

4 hours

S2 endurance
12,351 of 12,351 requests were accepted with no thermal failure or sustained queue accumulation.

Human-readable process

  1. Verify the GPU and test environment.
  2. Establish the strongest stable native baseline.
  3. Freeze S2 policies before comparison.
  4. Test native and S2 under matched conditions.
  5. Preserve passing, failing and invalid runs.
  6. Publish only evidence-supported conclusions.

Evidence-package preview

  • Test environment
  • Hardware and model classification
  • Native and S2 capacity boundaries
  • Quality and service-level outcomes
  • Physical-call reduction
  • GPU-board energy results
  • Failed and excluded runs
  • Integrity verification

The public package will not include raw decision traces, protected configurations, workload seeds or materials that reveal internal decision logic.

Validation status

Laptop validation and the A100 fixed-demand POC are complete within their stated boundaries.

01

Complete

Mechanisms validated

Accepted-response reuse, concurrent sharing, supersession, scheduling, thermal control and workload separation were exercised with retained evidence.

02

Complete

Cross-model effect confirmed

Five valid fixed-demand pairs across four qualified Ollama models all showed fewer physical calls and lower GPU-board energy.

03

Complete

Capacity and endurance confirmed

The Llama 3.1 laptop ceiling was directly bounded, repeated in six one-hour trials and sustained in a four-hour S2 endurance run.

04

Complete

A100 fixed-demand POC completed

Three valid matched pairs on a full A100 through vLLM showed a positive indicative paired effect with effectively equivalent accepted service.

05

Pending

Higher assurance remains

A100 capacity-ceiling testing, measured live negative controls, independent reproduction, production-cluster validation and whole-system energy measurement have not been completed.

Current A100 classification: self-run hosted-GPU fixed-demand POC using a full NVIDIA A100-SXM4-80GB allocation hosted by Google Colab, Qwen2.5-7B-Instruct BF16 and vLLM 0.23.0. The laptop and A100 evidence are not independent validation or production-cluster validation.

Customer whitepaper

Download the completed customer-facing S2 position.

The whitepaper summarises the platform, workload fit, completed A100 and laptop evidence, assessment path and explicit claim boundaries. A separate evidence request can still be prepared for the bounded A100 position.

04

Frequently asked questions

Clear answers before a confidential discussion.

The public site explains the purpose, boundary and evidence standard. It does not teach the protected implementation.

What is S2?

S2 is a software capacity-control platform placed between AI applications and model serving. It is designed to reduce avoidable physical model calls while preserving accepted service outcomes.

Who is S2 for?

S2 is for organisations operating AI inference workloads with observable capacity, latency or queue pressure.

What problem does it solve?

It addresses the gap between logical demand and necessary physical computation. Retries, repeated demand, simultaneous demand and obsolete work can consume GPU capacity without creating equal value.

What is a logical request?

A logical request is a service obligation presented by an application. It counts as useful service only when the delivered outcome meets the agreed requirements.

What is a physical model call?

A physical model call is a fresh execution request sent to the model-serving environment. S2 measures how many fresh calls are required to serve accepted logical demand.

What is quality-accepted service?

It is a delivered outcome that passes the quality and service requirements agreed for the workload. Rejected outcomes do not count as useful capacity.

Does S2 replace the existing model server?

No. S2 is positioned before the existing model-serving environment. Required work continues to the agreed model and serving platform.

Does S2 require model retraining?

The public S2 position does not require changing or retraining the model. The model obligation remains unchanged.

Where does S2 sit in the serving environment?

S2 sits at the boundary between application demand and model serving. It is evaluated as a controlled service adjacent to the existing inference platform.

What remains under the customer’s control?

The customer retains control of model choice, serving environment, workload and tenant boundaries, quality requirements, service objectives and deployment approval.

Does S2 change the model obligation?

No. The model obligation is the agreed model and service requirement that must still be fulfilled. Work that requires fresh model execution is forwarded.

Does it make the GPU faster?

No. S2 is intended to increase useful service from available capacity by reducing avoidable work, not by changing GPU performance.

What happens when optimisation is unavailable?

S2 fails open: required work is forwarded when an optimisation cannot be used safely or is unavailable.

Can S2 help a workload made entirely of unique requests?

Opportunity may be limited when almost every request is unique and necessary. The assessment may conclude that S2 is not the right next step.

How is an acceptable outcome defined?

The customer and assessment team agree workload-specific quality and service requirements before comparison. A result is useful only when it passes those requirements.

What does the A100 fixed-demand POC show?

In one self-run Google Colab session using a full A100-SXM4-80GB, Qwen2.5-7B-Instruct BF16 and vLLM 0.23.0, three valid matched pairs retained effectively equivalent quality-accepted service while Full S2 used 26.19% fewer physical model calls and approximately 1.76% less GPU-board energy per accepted request.

Does the A100 result prove a 35.48% capacity uplift?

No. The 35.48% figure is a calculated physical-work-equivalent indicator. The POC did not search for native or S2 capacity ceilings, admit additional logical demand or establish a measured capacity uplift.

What are the principal A100 limitations?

The result comes from one self-run Colab session and three deliberately controlled fixed-demand pairs. The protocol did not include measured live unique-work negative controls or sensitivity conditions, and its 20% retry and 10% fan-out traffic profile is not assumed to represent an individual customer.

What does the completed laptop evidence show?

Across five valid fixed-demand pairs covering four qualified Ollama models, S2 used 28.19% fewer physical calls and 16.10% less GPU-board energy per accepted request while serving equivalent accepted demand.

What capacity result has been established?

For Llama 3.1 8B Q4_K_M through Ollama 0.31.1 on the tested RTX 5050 laptop, S2 sustained at least 50.09% more quality-accepted logical workload than the confirmed native ceiling.

What energy boundary is measured?

The completed validation measures GPU-board energy at the GPU device boundary. It does not claim whole-laptop, whole-server, rack, cooling or facility energy savings.

Why did latency not improve in every comparison?

S2 is not presented as a universal latency reduction. The valid comparisons maintained quality and service attainment, but latency did not improve in every pair.

Which serving environments have been demonstrated?

The completed evidence contains separate strata: Ollama on an RTX 5050 laptop across four qualified models, and a fixed-demand POC using Qwen2.5-7B-Instruct BF16 through vLLM 0.23.0 on a full A100-SXM4-80GB hosted by Google Colab. The two strata are not pooled.

Is S2 production-ready?

The site does not claim production readiness. Production-cluster validation, independent reproduction and customer-specific operational assurance have not been completed.

Is the current validation independent?

No. Both the laptop programme and the hosted A100 POC are self-run. Independent reproduction and production-cluster validation have not been completed.

Has A100 or vLLM testing been completed?

Yes, within a narrow POC boundary. Three valid fixed-demand pairs were completed through vLLM 0.23.0 on a full Google Colab A100 allocation. This is indicative paired evidence, not a measured A100 capacity ceiling or production validation.

How can an organisation request an assessment?

Prepare a non-confidential request on the assessment page. The website opens a draft addressed to info@proggen.co.uk; nothing is sent until the visitor reviews and sends it through their email service.

What happens after an assessment request is sent?

ProgGen reviews the non-confidential information, discusses workload fit and, if further evaluation is justified, moves technical detail into an appropriate confidential engagement. A pilot is not guaranteed.

What information should not be sent by email?

Do not send confidential workload data, source code, credentials, raw prompts, protected configurations or detailed internal records in the initial message.

What could make S2 unsuitable?

A no-go result may follow when demand is almost entirely unique and necessary, acceptable output cannot be defined, pressure cannot be measured, boundaries are unclear or a controlled native comparison cannot be established.

Evidence before expansion

Define the boundary. Preserve the proof.

Request the completed A100 POC summary or begin with a workload-fit discussion.

Evidence includes

  • Accepted service outcomes
  • Physical-call reduction
  • Failed and excluded runs