Continuous tenant-isolation verification for GPU clusters.

We design secure multi-tenant GPU clusters and continuously prove that every tenant stays isolated, from the fabric to the firmware.

Now onboarding design partners.

Why a razorbill

A shared GPU cluster is a crowded colony.

Workloads for banks, governments and AI labs, side by side on the same cliff.

Every tenant should hold its own ledge.

What can one tenant see, reach or change about another?

A canary checks every boundary.

Every time a node is provisioned or recycled, and on a schedule.

tenant-a tenant-b tenant-c tenant-d canary
The problem

Shared GPU clusters run workloads for banks, governments and AI labs side by side.

Recent independent testing of GPU cloud providers found tenants able to:

Exposure

See other tenants' infrastructure

Boundary crossing

Read across tenant boundaries

Lateral access

Reach management networks

Often through default settings rather than sophisticated attacks.

Configurations drift. Nodes get recycled. A one-off audit is out of date the day after it ends.

Providers need continuous evidence, and their customers increasingly ask for it.

How it works

A canary tenant that runs from an ordinary customer's position.

Runs every time a node is provisioned or recycled, and on a schedule.

It asks one question

What can one tenant see, reach or change about another?

It checks every shared boundary.

07 boundaries

Fabric

InfiniBand partitions and keys, RoCE and Ethernet segmentation.

Management plane

BMC, IPMI and Redfish reachability.

DPUs and SmartNICs

Operating mode and host access.

Orchestration

Kubernetes and Slurm exposure, network policy, shared control planes.

Storage and monitoring

Per-tenant separation in the backend, not just the dashboard.

Node reuse

Disk wipe, DPU reflash and firmware state between tenants.

Software

Known-vulnerable drivers, container toolkits, runtimes and firmware.

→ when

Every time a node is provisioned or recycled, and on a schedule.

On failure

Alert with the exact fix

Failures raise an alert with the exact fix.

Then

Re-test

Every fix is followed by a re-test.

Every handoff

Isolation report

Each tenant handoff produces an isolation report the provider can share with that customer.

What we offer

From first design to every node handoff.

01 / Design

Secure cluster design

Isolation designed in from the start, from fabric partitioning to management-network separation, then verified once the cluster is live.

For organisations building shared GPU clusters
02 / Audit

Baseline isolation audit

A full, scoped assessment of an existing cluster, with findings, fixes and a re-test.

03 / Verify

Continuous verification

Ongoing checks on every node handoff, with per-tenant reports for your customers.

Vendor-neutral NVIDIA&AMD InfiniBand&Ethernet Kubernetes&Slurm
How we work

Testing on your terms.

  • No test runs without a signed, written scope.

  • Your engineers stay in the loop.

  • Findings go only to you. Nothing is ever published without your written consent.

  • Checks run inside your environment, so your data stays in-country.

About

Saad Khan

Founder

Alcora Labs was founded by Saad Khan. He has five years of experience in AWS and cloud infrastructure, a background in electrical engineering, and comes from Instec, a cybersecurity firm that has operated across Pakistan and the Gulf since the 1980s.

Named after the razorbill: a seabird that nests in crowded colonies, each pair holding its own ledge.

Prove every tenant stays isolated.

Now onboarding design partners.