Enterprise AI, optimised for value, control and scale.Discover AI Economics

Control plane

Manage the platform.

The Control Plane is the central command and governance layer for the entire T-Flux Ultra environment. It provides a single, auditable view of platform health, workload, performance, governance, security posture, and capacity and cost.

It answers one operational question: is T-Flux operating securely, efficiently, and within policy?

T-Flux Ultra Control Plane dashboard showing platform health, infrastructure, security, retention, cost, capacity, alerts, and quick actions.

One operating view

Control every aspect of a customer’s T-Flux Ultra environment.

For single-tenant T-Flux installations in a private cloud or on-premises, the Control Plane brings service status, infrastructure performance, policy enforcement, and governance evidence together for platform administrators, enterprise IT, security, risk, and T-Flux operators.

Core functions

Operate with visibility, control, and evidence.

01

Platform health and service status

See uptime, SLO attainment, end-to-end latency, error rates, queues, and the oldest outstanding jobs.

02

Workload and capacity management

Monitor concurrent sessions, storage growth, compute and GPU capacity, workload queues, and resource utilisation.

03

Security and governance

Bring security posture, access controls, governance policies, audit controls, alerts, incidents, and exceptions into one view.

04

Infrastructure performance

Review CPU, RAM, disk IOPS, network throughput, GPU utilisation, VRAM, temperature, and power across the environment.

05

Cost and resource utilisation

Track GPU-hours, CPU-hours, storage growth, cost proxies, cost per deliverable, and available headroom.

06

Configuration and policy enforcement

Manage platform configuration, administrative controls, cross-tenant monitoring, and policy enforcement.

Views by role

Information for the people who operate the environment.

Customer Admin

Service status, consumption, API keys, governance, and billing or cost proxies.

Operations

Incidents, queues, latency, infrastructure, logs, and remediation actions.

Security and Risk

Policy changes, access events, DLP events, audit integrity, and retention compliance.

Data Steward

Ingestion quality, re-indexing, retrieval health, and corpus change control.

Model Admin

Model routing, throughput, quality KPIs, GPU scheduling, and model failures.

Dashboard coverage

See the whole operating picture.

Operational tiles update within ten seconds and governance tiles within sixty seconds, giving teams current, auditable visibility across the platform.

A

Overview

Uptime across 24 hours, 7 days, and 30 days; SLO attainment; latency; requests; concurrent sessions; error budget; queues; and storage growth.

B

Compute and models

GPU and VRAM utilisation, CPU and memory, disk and network performance, tokens per second, routing distribution, and model concurrency.

C

Data plane

Connector throughput, parse and OCR success, embedding queues, vector index health, retrieval latency, cache hit rate, and ACL-filtered retrieval.

D

Quality and governance

Citation coverage, grounding and verification scores, low-evidence refusals, revision rate, corpus changes, policy changes, and guardrail rule hits.

E

Security and audit

Authentication failures, privileged actions, break-glass usage, DLP and PII events, egress controls, audit integrity, and retention compliance.

F

API and integrations

API usage, rate limiting, API-key lifecycle, webhook and callback health, retry counts, and integration latency.

G

Batches and agents

Batch jobs, duration distribution, workflow step timings, tool-call failures, retry rate, and poison-pill detection.

Board and audit-ready controls

Turn operating data into assurance and forward-looking management.

Security and audit

Surface privileged actions, break-glass use, policy changes outside workflow, DLP events, egress controls, and tamper-evident audit integrity.

Storage and retention

Track storage by artefact class, retention compliance, deletion failures, backup and restore status, RPO/RTO posture, and restore-test history.

Cost and capacity

Forecast GPU-hours, CPU-hours, storage growth, cost proxies, headroom, saturation, and the impact of capacity changes.

Alerts and thresholds

Make exceptions visible before they become business risk.

Default SEV1, SEV2, and SEV3 baselines cover availability, latency, error rate, queues, capacity, audit integrity, retention, and security. Every threshold can be configured per tenant and audited.

Availability

Below 99.5% daily is SEV2; below 99.0% is SEV1.

Latency and errors

p95 latency above target for a sustained period and elevated 5xx errors trigger escalation.

Queues and capacity

Queue depth, oldest-job age, GPU VRAM, disk utilisation, and vector index health are monitored against thresholds.

Security and retention

Audit-integrity failures, break-glass usage, unauthorised policy changes, and retention failures are surfaced for action.

Drill down and report

Trace the signal, investigate the cause, and export the evidence.

T-Flux Ultra Control Plane dashboard showing platform health, infrastructure, security, retention, cost, capacity, alerts, and quick actions.

Every widget can drill down to request traces, token statistics, model selection, retrieval hits, citations, job details, retries, infrastructure events, and correlated request, job, and document IDs.

One-click exports provide monthly SLA reports, governance and audit reports, and security posture summaries in PDF and CSV formats, each with timestamp, tenant, and signature hash.