Free

Cluster Triage

The free first step after a cluster incident. One read-only script across your Hyper-V, Failover Cluster or Azure Local environment. You receive the triage list: the five heaviest signals around your incident window, sorted by severity. Free, without subscription and without obligation.

WhatsApp · hans@cldlbs.com

TRIAGE · INCIDENT WINDOW
CRITICALWitness unreachable from both nodes
HIGHFirmware on NODE02 below the level the vendor prescribes
MediumPFC classes differ between host and switch

Example of the form; the content comes from your own measurement.

How it works

01

Request the script

Fill in the form and send it via WhatsApp and e-mail; you receive the script with its instructions.

02

Run it from your management server

It only reads: the event channels of all nodes over the incident window. It first shows the plan of the run, with one yes-or-no question. Memory dumps stay out of scope in this free variant.

03

Hand over the output

You place the file on the encrypted Proton share we set up for you.

You receive the triage list by e-mail.

Request

Five details are enough to send you the right script and instructions. The number of hours sets the measurement window of the run.

Send the request via both WhatsApp and e-mail, so we cannot miss it.

We aim to respond right away, subject to availability.
This form sends nothing by itself: the request leaves through your own WhatsApp or email. This page stores nothing.
WhatsApp · hans@cldlbs.com

What the triage list is, and is not

What you get

Five signals, sorted by severity, each with the measured value, the node and the timestamp. Whatever could not be read is listed too, with the reason.

What it is not

No cause and no reconstruction: triage sorts, the diagnosis takes deeper investigation. No report, no review session, and no subscription attached.

More than five found? The triage list also states the total number of signals found; the five heaviest are always in this free list.

Safe and read-only

  • Nothing is installed and nothing is changed; nothing is left behind afterwards.
  • No passwords are processed or stored. Before the file is written, the script checks its own output for values that look like credentials.
  • Memory dumps do not travel in the free variant.
  • What you hand over is used solely to produce the triage list.

If the list points at something

Triage sorts the signals, it does not explain them. When the list points at a real problem, the next step is a proper root-cause investigation: the reconstruction across all nodes, with timeline, gaps and an RCA document.

If the environment has never been measured end to end, that investigation is strongest combined with a baseline assessment through ClusterTriage Assurance: the same specialists, a read-only method, and a report built from measured state rather than a checklist.

Recognise this?

What Cluster Triage is for

These sentences come from findings written down in real engagements. If you recognise one, the triage list is the first step: it sorts the problem before you start searching.

Cluster and nodes

  • Cluster is down
  • Cluster node keeps failing
  • Node not added to the cluster
  • Cluster went down during patching
  • Cluster went down twice last weekend
  • Bluescreen on a cluster node
  • No crash dump after an outage
  • Event 1135: node removed from the cluster
  • Cluster validation reports errors

Quorum and storage

  • Quorum lost
  • Witness unreachable
  • Storage Spaces Direct volume degraded
  • S2D repair job does not finish
  • CSV went offline
  • Cluster Shared Volume in redirected mode
  • Event 5120: CSV I/O pause
  • High latency on storage
  • MPIO path lost

Network and migration

  • Live migration fails
  • Live migration is slow
  • RDMA does not work
  • Virtual machines unreachable after a restart
  • VM freezes randomly

Azure Local and backup

  • Azure Local update fails
  • Solution update hangs
  • Readiness check failed on Azure Local
  • Backup fails on the cluster
  • Veeam backup on the cluster is slow

If your situation is not listed, triage is still the safest first step: it only reads and changes nothing. What is needed after that emerges from the list.

Is your cluster down?

Do not wait until the evidence is gone.

Request the read-only triage script. You receive the five heaviest signals free of charge and know which next step makes sense.