Skip to content

Troubleshooting

This page collects common day-2 checks for a running appliance.

Common Failures

Symptom First checks
Flux Kustomization is False kubectl -n flux-system describe kustomization <name> and render the same path locally with kubectl kustomize.
HelmRelease is not ready kubectl -n flux-system describe helmrelease <name> and inspect chart values.
Custom legacy Ingress has no endpoint The nginx controller is intentionally not installed. Bundled surfaces already use Envoy; migrate custom applications to an authenticated HTTPRoute.
App waits for model catalog Check ai-model-catalog-controller logs and AI_APPLIANCE_MODEL_CATALOG_READY.
OpenClaw reports that auto-compaction cannot recover immediately in a new chat Check that openclaw.json contains the generated agents.defaults.compaction policy and tools.profile: coding, then restart the OpenClaw instance so its operator-managed configuration is reapplied. Do not raise the reserve floor to the full context size of a small local model.
Hermes instance URL returns 404: Not Found The agent gateway on port 8443 has no root UI. Confirm the Hermes instance enables HERMES_DASHBOARD=true, its Service exposes port 9119, and the generated HTTPRoute targets port 9119.
LiteLLM Prisma reports P1000 authentication failed The PostgreSQL PVC may be older than litellm-postgresql-secret. Keep generated DB credentials prune-disabled and rotate the DB user password to match the current Secret.
LiteLLM Models shows Virtual Key expected with a token beginning eyJ The Keycloak JWT replaced LiteLLM's own API key. Confirm both static-litellm-*-sso policies set spec.oidc.forwardAccessToken: false, reconcile app-litellm, and reload the LiteLLM UI.
An AI UI stops a longer answer with TypeError: network error, Could not respond, or An error occurred while streaming response Check the Envoy access log for response_code_details=response_timeout, response_flags=UT, and a duration near 15000 ms. Confirm the application's local/public HTTPRoute sets spec.rules[].timeouts.request: "0s". Reconcile the static module or magicstick-operator as appropriate.
Application shows a second login after SSO Confirm Paperclip uses deployment.mode: local_trusted with exposure: private, the patched operator emits PAPERCLIP_BIND=loopback, and its gateway-loopback-proxy sidecar is ready; confirm Odysseus has AUTH_ENABLED=false and the application Service is exposed only through its authenticated Envoy route.
Paperclip task reports that provider kubernetes is not registered or creates no Sandbox Check that plugin paperclip.kubernetes-sandbox-provider is ready, then check sandboxes.agents.x-k8s.io, the Agent Sandbox controller, PAPERCLIP_K8S_ADAPTER_TYPE=opencode_local, spec.adapters.execution.kubernetes.backend, and the selected adapter runtime image. For K3s, confirm control-plane egress TCP 6443; for Rancher namespace admission failures, confirm the instance-specific updatepsa ClusterRoleBinding.
Paperclip onboarding discovers OpenCode models but the hello probe times out Verify that the Paperclip server Pod can connect to litellm.ai.svc.cluster.local:4000 and that its additive NetworkPolicy permits TCP 4000 only to Pods labeled app=litellm. OpenCode retries an immediate connection refusal with backoff, so this otherwise appears as a slow 60-second probe.
Paperclip task stops after Sandbox run log streaming enabled Find the namespace labeled paperclip.io/managed-by=paperclip-k8s-plugin and verify NetworkPolicy/magicstick-paperclip-runtime-egress exists. The policy must select paperclip.io/role=agent and permit only the owning Paperclip server on TCP 3100 plus app=litellm on TCP 4000. Check magicstick-operator logs and RBAC if it is absent.
Paperclip task repeatedly runs for minutes and LiteLLM reports ContextWindowExceededError Check that paperclip-opencode-providers.json advertises a smaller context than the physical model limit. For the bundled 20000-token Qwen preset the Paperclip value is 15904 with output 3976. Reconcile the model catalog and restart the affected Paperclip run after the generated provider file changes.
Paperclip task spends most of its run searching instructions or reports skill ripgrep execution failed Confirm the Sandbox uses the pinned official Paperclip OpenCode runtime and that rg --version succeeds in it. Confirm the Paperclip Pod mounted the guarded adapter patch and its source-validation init container completed. Runs still keep the upstream 15-minute transport ceiling; repeated searches are a runtime or instruction-contract failure, not a reason to raise that timeout.
Paperclip run succeeds but generated files are absent from the task workspace, or logs show PAPERCLIP_WORKSPACE_CWD=/tmp while tools write below /workspace Confirm the guarded OpenCode adapter patch is mounted and the init container found its exact pinned source. Paperclip v2026.707.0 drops the Kubernetes realizeWorkspace result before transport resolution; Magic Stick normalizes only the Kubernetes /tmp fallback to the runtime's /workspace mount. Restart the Paperclip Pod after updating the template and use a fresh run/session.
Paperclip run creates local files but neither task documents nor a final issue status Confirm the agent has paperclipai/paperclip/paperclip in paperclipSkillSync.desiredSkills and its bootstrapPromptTemplate contains [MagicStick Paperclip heartbeat v2]. Availability alone does not force OpenCode to load a skill; the managed bootstrap directive makes the upstream Paperclip heartbeat procedure explicit without modifying the skill.
A sandboxed Paperclip agent reports that Paperclip is unavailable on localhost:3020 or localhost:3100 Those literal ports are invalid inside the run sandbox. Paperclip injects a run-scoped callback bridge through PAPERCLIP_API_URL; the Magic Stick heartbeat bootstrap directive requires every agent API request to use that value. Reconcile the instance and start a fresh agent session so the current directive is applied.
Paperclip run stays pending and the tenant namespace reports exceeded quota: paperclip-quota Compare each Sandbox paperclip.io/run-id with the Paperclip heartbeat run. The instance helper removes only terminal runs after a 60-second grace; inspect the gateway-loopback-proxy log if they remain. Confirm the instance-specific sandbox-reconciler ClusterRole adds only list; the operator-owned execution role already supplies Sandbox deletion. Do not increase the quota to mask orphaned runtime Pods.
Paperclip sandbox cannot call a model Check paperclip-opencode-providers.json, litellm-masterkey-secret, LiteLLM on port 4000, and NetworkPolicies in the Paperclip tenant namespace.
Generated Secret missing Check the secret generator HelmRelease and Secret annotations.
OIDC route does not redirect Check the SecurityPolicy and HTTPRoute status, Keycloak readiness, the Envoy data-plane logs, and whether the identity and requested application hostnames resolve to the Envoy LoadBalancer address.
AppInstance route returns 403 after SSO Compare spec.access.role with the user's magicstick-user, magicstick-viewer, magicstick-operator, or magicstick-admin realm roles.
Static AI route returns 403 after SSO AI routes require at least magicstick-user. Check the user's realm roles and the corresponding static SecurityPolicy.
Dashboard returns 403 after login Confirm the user has magicstick-viewer, magicstick-operator, or magicstick-admin; configuration changes need operator or admin as documented in authentication.md.
magicstick login cannot resolve api.magicstick.local Confirm dashboard-api-local is accepted, carries lab42.io/mdns.enabled: "true", the Gateway has an address, and kdns has published the route. Override MAGICSTICK_API_URL only for the actual appliance hostname.
magicstick login reports that no device endpoint is advertised Confirm the magicstick-cli client exists in the magicstick realm with Device Authorization Grant enabled, and let the Keycloak post-start reconciliation complete.
CLI/TUI reports a local certificate verification error Trust the public Magic Stick CA in the operating system or run the first login with --ca-file /path/to/magicstick-oidc-ca.crt. MAGICSTICK_CA_FILE and NODE_EXTRA_CA_CERTS are also supported. For a disposable system on a trusted test network only, use --insecure; it is process-local and prints a warning.
CLI receives 401 access token client is not trusted Confirm the dashboard API Deployment uses OIDC_EXPECTED_CLIENT_IDS=magicstick-human-gateway-local,magicstick-cli, then log in again so the token has azp=magicstick-cli.
System → Users is missing for an administrator Confirm the session contains magicstick-admin and the installation uses local Keycloak rather than the direct-external-provider escape hatch. Refresh the browser after role changes.
System → Users reports that Keycloak is unavailable Check Keycloak readiness, the dashboard API logs, the existence of magicstick-user-admin-client, and its exact-name Secret Role. Do not decode the Secret.
Federated SSO tab is missing or unavailable Confirm a live magicstick-admin, local Keycloak mode, the federated-sso entitlement, magicstick-federation-admin-client, and its exact-name Secret Role. Wait for the Keycloak post-start client reconciliation; never grant the general user-admin client identity-provider rights.
Federated metadata validation fails Confirm the discovery/metadata URL is HTTPS, reachable from Keycloak, presents a trusted TLS certificate, and contains the required OIDC endpoints or an active SAML IdP entity, SSO service and signing certificate. The dashboard intentionally rejects expired/unsigned SAML metadata, embedded URL credentials, HTTP endpoints and raw Keycloak fields.
A managed upstream login disappears after license change Check System → License. Missing, invalid or expired federated-sso entitlement disables managed providers fail closed; restore a valid entitlement, review the provider, enter the OIDC secret again, validate metadata and explicitly save it enabled. Local recovery login remains available.
User change returns 409 Check whether the account is external, protected, the current actor, or the last enabled local administrator. Duplicate username or email also returns 409.
API Access tab is missing Confirm the current session contains magicstick-admin, then refresh after any role change. Unlike Users, this tab does not depend on local Keycloak user-administration mode.
API Access reports LiteLLM key management unavailable Check the LiteLLM Pod and Service, PostgreSQL readiness, and the existence of ai/litellm-masterkey-secret without decoding it. Lost raw keys cannot be recovered; create a replacement and revoke the old named access.
Kubernetes Access tab is missing Confirm the current session contains magicstick-admin, local Keycloak identity management is active, and refresh after any role change.
Kubernetes kubeconfig download is disabled Check identity-system/magicstick-kubernetes-access-info. On appliance K3s, rerun the host converge so the identity CA is installed and the API server is restarted with OIDC. On an existing cluster, complete the platform-specific API-server configuration and publish the marker as documented.
kubectl oidc-login cannot open or complete login Install the kubelogin plugin, verify id.<mdns-domain> resolves from the workstation, trust only the CA embedded in the kubeconfig, and ensure loopback ports 8000 or 18000 are available.
OpenLens reports lookup <appliance>.local ... no such host Download the kubeconfig again. Appliance kubeconfigs use the current private control-plane IP for the Kubernetes API because the OpenLens proxy may bypass mDNS. If DHCP changed the address again, rerun host convergence or wait for the Ready node address to be visible, then download a fresh file. Keep the OIDC issuer on id.<mdns-domain>.
Kubernetes login succeeds but RBAC is denied Confirm the user has exactly one direct magicstick-kubernetes-* group, obtain a fresh token after the group change, and inspect the oidc: Group subjects in the ClusterRoleBindings.
GPU model never starts Check the vendor GPU operator, allocatable GPU resources, KubeAI model status, and the selected vLLM/Ollama server logs.
The physical monitor goes black on an NVIDIA-only appliance, but SSH and the setup page work Check cat /proc/fb, lsmod | grep nvidia_drm, cat /sys/module/nvidia_drm/parameters/{modeset,fbdev}, and systemctl status magicstick-setup-console.service. The NVIDIA display host role must install the pinned driver and load nvidia-drm with modeset/fbdev enabled; the K3s node must have nvidia.com/gpu.deploy.driver=false so the GPU Operator does not replace the console-owning module. A newly installed NVIDIA host receives one scheduled reboot after successful convergence. CPU-only and AMD-only hosts do not take this path. If the display remains black after reboot, use SSH for diagnosis; do not remove working setup state.
GPU hardware is present but provider is Unsupported Confirm node architecture/Kubernetes preflight first. Broad AMD/Intel PCI detection does not imply vendor support. For additional AMD hardware, use only a catalogued, explicitly acknowledged compatibility profile with fresh host checks and GPU registration. Engine validation is optional; unknown cards remain unsupported. Never force the vendor's NFD support label.
AMD HelmRelease times out waiting for DeviceConfig/default with NoMatchingNodes The current chart disables Helm's default operand and Magic Stick manages its DeviceConfig separately using appliance.magicstick.dev/amd-gpu-eligible. Check that the new Helm values/controller are deployed, then inspect System → Hardware, upstream support or explicit profile consent, current host evidence and device resources. Controller installation and usable GPU readiness are separate.
AMD host driver is ready but DRA reports an old device or checkpoint after reboot Deploy the reviewed recovery.2 image and matching controller. Recovery keeps the plugin registered, withholds unsafe allocations and waits for kubelet Unprepare. Legacy claims are migrated only after all reservations and Pod references disappear. Never clear the checkpoint or force-delete a still-used shared claim. See safe recovery.
Strix Halo remains unverified or has no render/KFD devices Run sudo magicstick-gpu-preflight --json, inspect the actual kernel/firmware/driver, and follow explicit host preparation. PCI matching does not fix missing kernel support. The preparation role accepts reviewed package pins and never reboots automatically.
Experimental AMD engine is unavailable although a GPU is allocatable Engine validation is no longer a gate. Confirm the host evidence timer, Node UID/boot/kernel identity, profile consent and engine eligibility labels under System → Hardware. Model scheduling also waits for KubeAI to adopt the configured runtime image. Neither failed nor missing smoke tests disable a ready GPU.
An optional GPU engine test failed or became stale Inspect its Job logs and the actual model runtime. GPU use remains enabled when hardware is ready. Run GPU validation, CLI hardware validate --yes, or the TUI validation action starts a fresh bounded test; host preparation and reboots do not automatically repeat it. A passed test is not a model-size or memory-accounting guarantee.
Provider is Conflict A vendor CRD already existed without a Magic Stick ModuleActivation. Decide which installation owns the operator; do not run a second copy.
Provider remains Installing with zero resources The controller chart is installed but driver/device-plugin readiness is incomplete. Inspect the vendor namespace and the node's allocatable extended resources. AMD's baseline expects a working host/inbox amdgpu driver.
NVIDIA remains Installing although nvidia.com/gpu is allocatable The driver and device plugin are active, but the dashboard cannot read DCGM telemetry yet. Check the nvidia-dcgm-exporter Pod, Service, endpoints, and the dashboard API ServiceAccount's services/proxy permission.
Provider changes to Unknown after reboot NFD has temporarily lost the PCI signal. Magic Stick intentionally retains the existing operator; wait for the next 60-second NFD pass and inspect the node before taking action.
Local model stays in WaitingForGPU The optional runtime is installed but Kubernetes reports no allocatable nvidia.com/gpu; verify supported hardware, driver pods, and node capacity.
Accelerator target is disabled in the dashboard The matching vendor module must be Ready and a Ready schedulable node must expose nvidia.com/gpu, amd.com/gpu, gpu.intel.com/i915, or gpu.intel.com/xe. CPU remains available independently.
A Compute memory gauge shows metrics unavailable CPU first checks the Kubelet node summary and then metrics.k8s.io; verify the dashboard API ServiceAccount can read nodes/proxy and node metrics. NVIDIA requires DCGM. Dedicated AMD/Intel usage remains unknown without a compatible exporter. An explicit unified-memory AMD profile can expose host mapping bounds as shared capacity, not as measured GPU usage or extra RAM.
Intel model stays in WaitingForGPU Confirm whether the node publishes gpu.intel.com/xe or gpu.intel.com/i915; the resolved profile in ModelActivation.status must match that resource.
A GPU appears in Compute memory but not in the model hardware selector Confirm the vendor module is enabled and inspect the node's allocatable resource (nvidia.com/gpu, amd.com/gpu, gpu.intel.com/xe, or gpu.intel.com/i915). A transient Flux Reconciling phase no longer blocks selection once that resource exists; without it, the driver or device plugin is not ready.
CPU model stays in Starting Check the CPU model Pod for image-pull, RAM, CPU, model-download, or vLLM startup failures; no NVIDIA checks should appear.
CPU vLLM reports insufficient memory for KV cache Inspect the info overlay's cache formula and recreate through the current dashboard so spec.local.kvCacheMemoryBytes is derived from architecture, context, and sequences. Older/direct resources without it retain the 512 MiB fallback. Reduce context/concurrency or increase RAM. Explicit local.allowMemoryRisk: true permits a below-estimate trial, but does not shrink the derived cache or guarantee startup.
CPU vLLM is OOM-killed while loading or warming a quantized or multimodal model The current estimator includes checkpoint bytes, a possible runtime working-weight copy, compile/warm-up headroom, the multimodal processor cache, and a conservative hybrid-cache factor. Confirm that the ModelActivation contains both memoryRequiredMi and kvCacheMemoryBytes, then compare the Pod limit and cgroup peak. If the recommendation is larger than the node, reduce context or choose a smaller model instead of raising only the timeout.
KV cache remains pending confirmation Compare status.requestedKvCacheType with the empty status.effectiveKvCacheType, then inspect the generated KubeAI Model args/env and model-pod logs. The operator publishes the effective value only after the configured runtime has a Ready replica; it never labels a rejected or silently changed mode as active.
Hugging Face model search is unavailable or rate-limited Retry after the short-lived discovery cache can refresh, narrow a broad prefix, and inspect the dashboard API log without printing credentials. Model discovery uses only the public Hugging Face API. Tested presets and direct hf:// references remain available and do not depend on the search endpoint.
Ollama Library search or tag lookup is unavailable Retry after the short-lived discovery cache can refresh and inspect the dashboard API log. Discovery reads only bounded public ollama.com pages because Ollama does not document a remote catalog API. Tested presets and direct ollama:// references remain available; model creation is not coupled to Library discovery.
A selected Hugging Face quantization fails during model loading Treat dynamic community artifacts as experimental. Confirm the repository format and quantization are supported by the selected vLLM image and CPU/GPU generation, compare the memory estimate with the actual node, and try the original repository or a tested preset. Magic Stick discovers artifacts; it does not convert or validate every third-party quantization.
Local model stays in Starting Compare kubectl -n ai get model <name> -o jsonpath='{.status.replicas}' with the model pod readiness and selected engine logs. The model is intentionally absent from LiteLLM until at least one replica is ready.
Model is WaitingForPod or Degraded / ModelPodCreationStalled with no Pod Inspect kubectl -n ai logs deployment/kubeai and the API-server/K3s journal for Pod-create or admission errors. The controller reports a stall after two minutes without a non-terminating Pod and keeps retrying. status.podCreation persists the timer for the current model/configuration; downloads in an existing Pod do not trigger this timeout. For AMD DRA, verify magicstick-amd-dra admission uses plain JSON maps rather than typed CEL objects in claim arrays, then verify a server-side dry run under the KubeAI ServiceAccount.
Ollama model stays in Starting with WaitingForOllamaAlias The source tag is still downloading, the model API is unreachable, or the KubeAI-name alias is absent. Check ollama list in the model pod and the ModelActivation.status.message. The operator creates the alias automatically as soon as the source tag is complete; do not publish a manual LiteLLM route around this guard.
LiteLLM returns 404 model '<name>' not found for an Ollama model Confirm the ModelActivation is Ready under the current operator revision. Older revisions trusted KubeAI replica readiness before the Ollama alias existed. Reconcile or restart the Magic Stick Operator after updating; it verifies the alias through /api/tags and repairs it through /api/copy.