Custom Resource Guide
The Splunk Operator provides a collection of custom resources you can use to manage Splunk Enterprise deployments in your Kubernetes cluster.
- Custom Resource Guide
- Metadata Parameters
- Common Spec Parameters for All Resources
- Common Spec Parameters for Splunk Enterprise Resources
- LicenseManager Resource Spec Parameters
- Standalone Resource Spec Parameters
- SearchHeadCluster Resource Spec Parameters
- Queue Resource Spec Parameters
- ClusterManager Resource Spec Parameters
- IndexerCluster Resource Spec Parameters
- IngestorCluster Resource Spec Parameters
- ObjectStorage Resource Spec Parameters
- MonitoringConsole Resource Spec Parameters
- Examples of Guaranteed and Burstable QoS
- Status Conditions
- Troubleshooting
For examples on how to use these custom resources, please see Configuring Splunk Enterprise Deployments.
Metadata Parameters
All resources in Kubernetes include a metadata section. You can use this to define a name for a specific instance of the resource, and which namespace you would like the resource to reside within:
| Key | Type | Description |
|---|---|---|
| name | string | Each instance of your resource is distinguished using this name. |
| namespace | string | Your instance will be created within this namespace. You must ensure that this namespace exists beforehand. |
If you do not provide a namespace, you current context will be used.
apiVersion: enterprise.splunk.com/v4
kind: Standalone
metadata:
name: s1
namespace: splunk
finalizers:
- enterprise.splunk.com/delete-pvc
The enterprise.splunk.com/delete-pvc finalizer is optional, and may be used to tell the Splunk Operator that you would like it to remove all the Persistent Volumes associated with the instance when you delete it.
Common Spec Parameters for All Resources
apiVersion: enterprise.splunk.com/v4
kind: Standalone
metadata:
annotations:
service.beta.kubernetes.io/azure-load-balancer-internal: "true"
name: example
spec:
disableResourceDefaults: false
imagePullPolicy: Always
livenessInitialDelaySeconds: 400
readinessInitialDelaySeconds: 390
podAnnotations:
traffic.sidecar.istio.io/excludeOutboundPorts: "8089,8191,9997,15020"
traffic.sidecar.istio.io/includeInboundPorts: "8000,8088,15021"
serviceTemplate:
spec:
type: LoadBalancer
topologySpreadConstraints:
- maxSkew: 1
topologyKey: zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
foo: bar
extraEnv:
- name: ADDITIONAL_ENV_VAR_1
value: "test_value_1"
- name: ADDITIONAL_ENV_VAR_2
value: "test_value_2"
resources:
requests:
memory: "512Mi"
cpu: "0.1"
limits:
memory: "8Gi"
cpu: "4"
The spec section is used to define the desired state for a resource. All custom resources provided by the Splunk Operator include the following configuration parameters:
| Key | Type | Description |
|---|---|---|
| image | string | Container image to use for pod instances (overrides RELATED_IMAGE_SPLUNK_ENTERPRISE environment variable |
| imagePullPolicy | string | Sets pull policy for all images (either “Always” or the default: “IfNotPresent”) |
| disableResourceDefaults | boolean | Prevents the operator from filling in default CPU and memory requests and limits. Defaults to false. Set to true to preserve resources exactly as provided, including an empty value. |
| livenessInitialDelaySeconds | number | Sets the initialDelaySeconds for Liveness probe (default: 300) |
| readinessInitialDelaySeconds | number | Sets the initialDelaySeconds for Readiness probe (default: 10) |
| extraEnv | EnvVar | Sets the extra environment variables to be passed to the Splunk instance containers. WARNING: Setting environment variables used by Splunk or Ansible will affect Splunk installation and operation |
| schedulerName | string | Name of Scheduler to use for pod placement (defaults to “default-scheduler”) |
| podAnnotations | map[string]string | Sets annotations on Splunk instance pods. These annotations can override operator-provided pod annotations, including the Istio sidecar traffic annotations. |
| affinity | Affinity | Kubernetes Affinity rules that control how pods are assigned to particular nodes |
| resources | ResourceRequirements | The settings for allocating compute resource requirements to use for each pod instance. Missing CPU and memory keys receive operator defaults unless disableResourceDefaults is true. The default settings should be considered for demo/test purposes. Please see Hardware Resource Requirements for production values. |
| serviceTemplate | Service | Template used to create Kubernetes Services |
| topologySpreadConstraint | TopologySpreadConstraint | Template used to create Kubernetes TopologySpreadConstraint |
Postgres Node Sidecar
Starting with Splunk Enterprise 10.6, SOK disables the Postgres node sidecar by default on all managed pods by injecting SPLUNK_NODE_SIDECAR_POSTGRES_DISABLED=true. No SOK-supported use case requires the sidecar on Kubernetes (DMX, SPL2, and Data Orchestrator are unsupported on CMP-K), and on 10.4+ it would otherwise silently start a database accumulating data in an unsupported configuration.
To re-enable the sidecar, set the variable to a different value in spec.extraEnv:
spec:
extraEnv:
- name: SPLUNK_NODE_SIDECAR_POSTGRES_DISABLED
value: "false"
Note: This override may be removed in a future SOK release. Enabling the Postgres sidecar in SOK deployments is not officially supported.
KV Store Default Type
SOK configures Splunk Enterprise pods to use local KV Store by default by injecting SPLUNK_KVSTORE_DEFAULT_TYPE=local into managed pods, except for IngestorCluster. Splunk Ansible consumes this variable and writes [kvstore] defaultKVStoreType to server.conf. This setting requires Splunk Enterprise 10.6.0 or later.
The only supported value is local. If the variable is set in spec.extraEnv, it must use the same value:
spec:
extraEnv:
- name: SPLUNK_KVSTORE_DEFAULT_TYPE
value: "local"
Common Spec Parameters for Splunk Enterprise Resources
apiVersion: enterprise.splunk.com/v4
kind: Standalone
metadata:
name: example
spec:
etcVolumeStorageConfig:
storageClassName: gp2
storageCapacity: 15Gi
varVolumeStorageConfig:
storageClassName: customStorageClass
storageCapacity: 25Gi
volumes:
- name: licenses
configMap:
name: splunk-licenses
licenseManagerRef:
name: example
clusterManagerRef:
name: example
serviceAccount: custom-serviceaccount
The following additional configuration parameters may be used for all Splunk Enterprise resources, including: Standalone, LicenseManager, SearchHeadCluster, ClusterManager, IndexerCluster and IngestorCluster:
| Key | Type | Description |
|---|---|---|
| etcVolumeStorageConfig | StorageClassSpec | Storage class spec for Splunk etc volume as described in StorageClass |
| varVolumeStorageConfig | StorageClassSpec | Storage class spec for Splunk var volume as described in StorageClass |
| volumes | Volume | List of one or more Kubernetes volumes. These will be mounted in all container pods as as /mnt/<name> |
| defaults | string | Inline map of default.yml overrides used to initialize the environment |
| defaultsUrl | string | Full path or URL for one or more default.yml files, separated by commas |
| licenseUrl | string | Full path or URL for a Splunk Enterprise license file |
| licenseManagerRef | ObjectReference | Reference to a Splunk Operator managed LicenseManager instance (via name and optionally namespace) to use for licensing |
| clusterManagerRef | ObjectReference | Reference to a Splunk Operator managed ClusterManager instance (via name and optionally namespace) to use for indexing |
| monitoringConsoleRef | string | Logical name assigned to the Monitoring Console pod. You can set the name before or after the MC pod creation. |
| serviceAccount | ServiceAccount | Represents the service account used by the pods deployed by the CRD |
| extraEnv | Extra environment variables | Extra environment variables to be passed to the Splunk instance containers |
| readinessInitialDelaySeconds | readinessProbe initialDelaySeconds | Defines initialDelaySeconds for Readiness probe |
| livenessInitialDelaySeconds | livenessProbe initialDelaySeconds | Defines initialDelaySeconds for the Liveness probe |
| imagePullSecrets | imagePullSecrets | Config to pull images from private registry. Use in conjunction with image config from common spec |
| certs | []CertSpec | List of TLS certificates to mount into Splunk pods. Each entry references a Kubernetes Secret containing tls.crt, tls.key, and optionally ca.crt. An optional role field (server or input) controls the mount path: server mounts at /mnt/tls/splunk-server-tls-cert/, input at /mnt/tls/splunk-input-tls-cert/, and no role mounts at /mnt/tls/<secretName>/. When a referenced Secret’s content changes, the operator automatically triggers a rolling restart. Note: cert secret rotation detection is supported for v4 CR types only (Standalone, LicenseManager, SearchHeadCluster, ClusterManager, IndexerCluster, MonitoringConsole, IngestorCluster). The deprecated v3 types (ClusterMaster, LicenseMaster) carry the certs field but do not watch for Secret changes. |
LicenseManager Resource Spec Parameters
apiVersion: enterprise.splunk.com/v4
kind: LicenseManager
metadata:
name: example
spec:
volumes:
- name: licenses
configMap:
name: splunk-licenses
licenseUrl: /mnt/licenses/enterprise.lic
Please see Common Spec Parameters for All Resources and Common Spec Parameters for All Splunk Enterprise Resources. The LicenseManager resource does not provide any additional configuration parameters.
Standalone Resource Spec Parameters
apiVersion: enterprise.splunk.com/v4
kind: Standalone
metadata:
name: standalone
labels:
app: SplunkStandAlone
type: Splunk
finalizers:
- enterprise.splunk.com/delete-pvc
In addition to Common Spec Parameters for All Resources and Common Spec Parameters for All Splunk Enterprise Resources, the Standalone resource provides the following Spec configuration parameters:
| Key | Type | Description |
|---|---|---|
| replicas | integer | The number of standalone replicas (miminum of 1, which is the default) |
SearchHeadCluster Resource Spec Parameters
apiVersion: enterprise.splunk.com/v4
kind: SearchHeadCluster
metadata:
name: example
spec:
replicas: 5
In addition to Common Spec Parameters for All Resources and Common Spec Parameters for All Splunk Enterprise Resources, the SearchHeadCluster resource provides the following Spec configuration parameters:
| Key | Type | Description |
|---|---|---|
| replicas | integer | The number of search heads cluster members (minimum of 3, which is the default) |
Search Head Deployer Resource
Since Search Head Deployer doesn’t require as many resources as Search Head Peers themselves, then Splunk Operator for Kubernetes 2.7.1 introduced additional field for SearchHeadCluster spec to manage resources for the deployer separately.
If provided, resources are managed separately for Search Head Deployer and Search Head Peers. Otherwise, either default values are used if resources are not defined at all or Search Head Peers resources are applied to Search Head Deployer as well.
Additionally, node affinity specification was introduced for Search Head Deployer to separate it from Search Head Peers specification.
| Key | Type | Description |
|---|---|---|
| deployerNodeAffinity | *corev1.NodeAffinity | Search Head Deployer node affinity |
| deployerResourceSpec | corev1.ResourceRequirements | Search Head Deployer resource specification |
Example
deployerNodeAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
...
requiredDuringSchedulingIgnoredDuringExecution:
...
deployerResourceSpec:
claims:
...
limits:
...
requests:
...
apiVersion: enterprise.splunk.com/v4
kind: SearchHeadCluster
metadata:
name: shc
finalizers:
- enterprise.splunk.com/delete-pvc
spec:
image: splunk/splunk: 9.4.4
serviceAccount: splunk-service-account
resources:
requests:
memory: "1024Mi"
cpu: "0.2"
limits:
memory: "10Gi"
cpu: "6"
deployerResourceSpec:
requests:
memory: "512Mi"
cpu: "0.1"
limits:
memory: "8Gi"
cpu: "4"
Queue Resource Spec Parameters
apiVersion: enterprise.splunk.com/v4
kind: Queue
metadata:
name: queue
spec:
provider: sqs
sqs:
name: sqs-test
authRegion: us-west-2
endpoint: https://sqs.us-west-2.amazonaws.com
dlq: sqs-dlq-test
To use static AWS credentials instead of workload identity, add the following optional secretKeyRef under spec.sqs:
secretKeyRef:
awsAccessKey:
name: s3-secret
key: s3_access_key
awsSecretKey:
name: s3-secret
key: s3_secret_key
Queue stores the configuration for an external message queue and dead-letter queue. SOK does not create or manage those external resources. Queue inputs can be found in the table below. The supported provider is sqs.
| Key | Type | Required | Description |
|---|---|---|---|
| provider | string | Yes | Provider of message queue (Allowed value: sqs) |
| sqs | SQS | Yes if provider = sqs | SQS message queue inputs |
SQS message queue inputs can be found in the table below.
| Key | Type | Required | Description |
|---|---|---|---|
| name | string | Yes | Name of the physical queue |
| authRegion | string | No | Region used for authentication and endpoint resolution |
| endpoint | string | No | AWS SQS service endpoint. If omitted, SOK resolves it from authRegion |
| dlq | string | Yes | Name of the physical dead-letter queue |
| secretKeyRef | object | No | Per-key selectors for AWS credentials. When not set, IRSA / workload identity is assumed. Contains awsAccessKey and awsSecretKey, each a SecretKeySelector with name and key. |
The provider, queue name, auth region, endpoint, and dead-letter queue are immutable after creation. secretKeyRef can be changed. If static credentials are configured, the referenced Secrets must be kept in the same namespace as the resource that uses them. SOK resolves the selected Secret keys and mounts generated credential-only defaults into the referenced IndexerCluster and IngestorCluster pods. Changes to the referenced credential Secret are watched. SOK creates a new credential Secret and rolls the affected pods declaratively.
ClusterManager Resource Spec Parameters
ClusterManager resource does not have a required spec parameter, but to configure SmartStore, you can specify indexes and volume configuration as below -
apiVersion: enterprise.splunk.com/v4
kind: ClusterManager
metadata:
name: example-cm
spec:
smartstore:
defaults:
volumeName: msos_s2s3_vol
indexes:
- name: salesdata1
remotePath: $_index_name
volumeName: msos_s2s3_vol
- name: salesdata2
remotePath: $_index_name
volumeName: msos_s2s3_vol
- name: salesdata3
remotePath: $_index_name
volumeName: msos_s2s3_vol
volumes:
- name: msos_s2s3_vol
path: <remote path>
endpoint: <remote endpoint>
secretRef: s3-secret
IndexerCluster Resource Spec Parameters
apiVersion: enterprise.splunk.com/v4
kind: IndexerCluster
metadata:
name: example
spec:
replicas: 3
clusterManagerRef:
name: example-cm
Note: clusterManagerRef is required field in case of IndexerCluster resource since it will be used to connect the IndexerCluster to ClusterManager resource.
In addition to Common Spec Parameters for All Resources and Common Spec Parameters for All Splunk Enterprise Resources, the IndexerCluster resource provides the following Spec configuration parameters:
| Key | Type | Required | Description |
|---|---|---|---|
| replicas | integer | Yes | The number of indexer peers. Must be at least 3 |
| queueRef | corev1.ObjectReference | No | Message queue reference. Set together with objectStorageRef to enable index-only mode |
| objectStorageRef | corev1.ObjectReference | No | Object storage reference. Set together with queueRef |
When both references are set, SOK configures the indexer peers to consume from the remote queue and use the object storage for large messages.
For reference update behavior that also applies to IndexerCluster, see Queue and ObjectStorage Reference Updates.
IngestorCluster Resource Spec Parameters
apiVersion: enterprise.splunk.com/v4
kind: IngestorCluster
metadata:
name: ic
spec:
replicas: 3
queueRef:
name: queue
objectStorageRef:
name: os
Note: queueRef and objectStorageRef are required fields in case of IngestorCluster resource since they will be used to connect the IngestorCluster to Queue and ObjectStorage resources.
In addition to Common Spec Parameters for All Resources and Common Spec Parameters for All Splunk Enterprise Resources, the IngestorCluster resource provides the following Spec configuration parameters:
| Key | Type | Required | Description |
|---|---|---|---|
| replicas | integer | No | The number of ingestor pods (defaults to 1) |
| queueRef | corev1.ObjectReference | Yes | Message queue reference |
| objectStorageRef | corev1.ObjectReference | Yes | Object storage reference |
Queue and ObjectStorage Reference Updates
Although the Queue and ObjectStorage configuration values are immutable after creation, these references can be changed. Changing either reference causes SOK to regenerate the content-addressed defaults resources and update the corresponding StatefulSet declaratively.
There is no supported migration strategy for moving data from previously referenced resources, which means that the existing data will not be available through the new configuration.
ObjectStorage Resource Spec Parameters
apiVersion: enterprise.splunk.com/v4
kind: ObjectStorage
metadata:
name: os
spec:
provider: s3
s3:
path: ingestion/smartbus-test
endpoint: https://s3.us-west-2.amazonaws.com
ObjectStorage stores the large messages that exceed the queue message-size limit. SOK does not create or manage the external bucket. ObjectStorage inputs can be found in the table below. The supported provider is s3.
| Key | Type | Required | Description |
|---|---|---|---|
| provider | string | Yes | Provider of object storage (Allowed value: s3) |
| s3 | S3 | Yes if provider = s3 | S3 object storage inputs |
S3 object storage inputs can be found in the table below.
| Key | Type | Required | Description |
|---|---|---|---|
| path | string | Yes | Remote storage location for messages that are larger than the underlying maximum message size |
| endpoint | string | No | S3-compatible service endpoint. If omitted, SOK resolves it from the Queue authRegion |
| encryptionScheme | string | No | Encryption scheme used by remote storage. Allowed values: sse-s3, sse-c, none |
| kmsEndpoint | string | No | KMS endpoint for generating data keys; auto-derived from the Queue region when not provided |
| kmsKeyId | string | No | ID of the primary KMS key (UUID, alias, or ARN) |
All ObjectStorage spec inputs are immutable after creation. kmsKeyId is required when encryptionScheme is sse-c.
MonitoringConsole Resource Spec Parameters
cat <<EOF | kubectl apply -n splunk-operator -f -
apiVersion: enterprise.splunk.com/v4
kind: MonitoringConsole
metadata:
name: example-mc
finalizers:
- enterprise.splunk.com/delete-pvc
EOF
Use the Monitoring Console to view detailed topology and performance information about your Splunk Enterprise deployment. See What can the Monitoring Console do? in the Splunk Enterprise documentation.
The Splunk Operator now includes a CRD for the Monitoring Console (MC). This offers a number of advantages available to other CR’s, including: customizable resource allocation, app management, and license management.
- An MC pod is not created automatically in the default namespace when using other Splunk Operator CR’s.
- When upgrading to the latest Splunk Operator, any previously automated MC pods will be deleted.
- To associate a new MC pod with an existing CR, you must update any CR’s and add the
monitoringConsoleRefparameter.
The MC pod is referenced by using the monitoringConsoleRef parameter. There is no preferred order when running an MC pod; you can start the pod before or after the other CR’s in the namespace. When a pod that references the monitoringConsoleRef parameter is created or deleted, the MC pod will automatically update itself and create or remove connections to those pods.
Examples of Guaranteed and Burstable QoS
You can change the CPU and memory resources, and assign different Quality of Services (QoS) classes to your pods. Here are some examples:
A Guaranteed QoS Class example:
Set equal requests and limits values for CPU and memory to establish a QoS class of Guaranteed.
Note: A pod will not start on a node that cannot meet the CPU and memory requests values.
Example: The minimum resource requirements for a Standalone Splunk Enterprise instance in production are 24 vCPU and 12GB RAM.
apiVersion: enterprise.splunk.com/v4
kind: Standalone
metadata:
name: example
spec:
imagePullPolicy: Always
resources:
requests:
memory: "12Gi"
cpu: "24"
limits:
memory: "12Gi"
cpu: "24"
A Burstable QoS Class example:
Set the requests value for CPU and memory lower than the limits value to establish a QoS class of Burstable.
Example: This Standalone Splunk Enterprise instance should start with minimal indexing and search capacity, but will be allowed to scale up if Kubernetes is able to allocate additional CPU and Memory up to the limits values.
apiVersion: enterprise.splunk.com/v4
kind: Standalone
metadata:
name: example
spec:
imagePullPolicy: Always
resources:
requests:
memory: "2Gi"
cpu: "4"
limits:
memory: "12Gi"
cpu: "24"
A BestEffort QoS Class example:
Set disableResourceDefaults to true and omit requests and limits to prevent the operator from adding resource defaults:
apiVersion: enterprise.splunk.com/v4
kind: Standalone
metadata:
name: example
spec:
disableResourceDefaults: true
resources: {}
With no requests or limits set for any container, Kubernetes assigns the pod the BestEffort QoS class. BestEffort QoS is not recommended for Splunk Enterprise production workloads. A namespace LimitRange may add resources during pod admission and result in a different QoS class.
Pod Resources Management
CPU Throttling
Kubernetes starts throttling CPUs if a pod’s demand for CPU exceeds the value set in the limits parameter. If your nodes have extra CPU resources available, leaving the limits value unset will allow the pods to utilize more CPUs.
Status Conditions
All Splunk Enterprise Custom Resources include Kubernetes-standard status conditions that provide detailed information about the resource state. These conditions follow Kubernetes conventions and can be used for monitoring, alerting, and automation.
Condition Types
| Condition Type | Description |
|---|---|
Ready | Indicates whether the resource is fully operational and all replicas are ready |
Progressing | Indicates whether the resource is being updated, scaled, or initialized |
Paused | Indicates whether reconciliation is paused via the pause annotation |
Stalled | Indicates a non-recoverable failure that requires user intervention before reconciliation can resume |
Condition Fields
Each condition includes the following fields:
| Field | Description |
|---|---|
type | The condition type (Ready, Progressing, Paused, or Stalled) |
status | Either “True”, “False”, or “Unknown” |
reason | A machine-readable reason code for the condition’s state |
message | A human-readable description of the condition |
lastTransitionTime | The last time the condition status changed |
observedGeneration | The generation of the CR spec that was observed |
Example Status with Conditions
status:
phase: Ready
conditions:
- type: Ready
status: "True"
reason: ReconcileComplete
message: Resource is ready
lastTransitionTime: "2026-05-04T10:00:00Z"
observedGeneration: 3
- type: Progressing
status: "False"
reason: Stable
message: Resource is stable
lastTransitionTime: "2026-05-04T09:55:00Z"
observedGeneration: 3
- type: Paused
status: "False"
reason: NotPaused
message: Reconciliation is not paused
lastTransitionTime: "2026-05-04T08:00:00Z"
observedGeneration: 3
- type: Stalled
status: "False"
reason: NotStalled
message: ""
lastTransitionTime: "2026-05-04T08:00:00Z"
observedGeneration: 3
When a terminal failure is detected, Stalled flips to True:
status:
phase: Error
conditions:
- type: Ready
status: "False"
reason: ReconcileFailed
message: Pod stuck in terminal state — manual fix required
lastTransitionTime: "2026-05-04T11:00:00Z"
observedGeneration: 4
- type: Progressing
status: "False"
reason: ReconcileFailed
message: Pod stuck in terminal state — manual fix required
lastTransitionTime: "2026-05-04T11:00:00Z"
observedGeneration: 4
- type: Paused
status: "False"
reason: NotPaused
message: Reconciliation is not paused
lastTransitionTime: "2026-05-04T08:00:00Z"
observedGeneration: 4
- type: Stalled
status: "True"
reason: PodTerminalFailure
message: Pod stuck in terminal state — manual fix required
lastTransitionTime: "2026-05-04T11:00:00Z"
observedGeneration: 4
Checking Conditions
You can view conditions using kubectl:
kubectl get standalone example -o jsonpath='{.status.conditions}' | jq .
Or describe the resource:
kubectl describe standalone example
Condition Behavior
lastTransitionTimeonly updates when the condition’sstatusfield changes (e.g., from “False” to “True”), not on every reconcileobservedGenerationreflects which spec generation the controller has processed- When an error occurs, the
Readycondition’smessagefield contains the specific error description Stalled=Truesignals a non-recoverable failure: the operator has stopped requeueing the CR and will not retry until the user resolves the root cause.Stalledis alwaysFalsewhenphaseis notError—Ready=TrueandStalled=Truecan never coexist- Use
Stalled=Truein monitoring or alerting rules to page on failures that need human intervention, as opposed to transient errors that self-heal - A Warning event with reason
Stalledis emitted on every reconcile whereStalled=True(not only on the initial flip); a Normal event with reasonStalledResolvedis emitted once when the condition clears fromTruetoFalse. Both are visible viakubectl describe
Troubleshooting
CR Status Message
The Splunk Enterprise CRDs with the Splunk Operator have a field cr.Status.message which provides a detailed view of the CR’s current status.
Here is an example of a Standalone with a message indicating an invalid CR config:
bash% kubectl get stdaln
NAME PHASE DESIRED READY AGE MESSAGE
ido Error 0 0 26s invalid Volume Name for App Source: custom. volume: csh, doesn't exist
bash# kubectl get stdaln -o yaml | grep -i message -A 5 -B 5
appsStatusMaxConcurrentAppDownloads: 5
bundlePushStatus: {}
isDeploymentInProgress: false
lastAppInfoCheckTime: 0
version: 0
message: 'invalid Volume Name for App Source: custom. volume: csh, doesn''t exist'
phase: Error
readyReplicas: 0
replicas: 0
resourceRevMap: {}
selector: ""
Terminal Failures
Some failure states are non-recoverable without external intervention. When the operator detects one, it stops reconciling the CR immediately — the CR is not requeued — and sets status.phase to Error with Stalled=True in the status conditions. The CR remains in this state until the root cause is resolved and the operator detects the change.
What triggers a terminal failure
| Cause | Stalled condition message | Affected CRs |
|---|---|---|
A container is stuck in a non-recoverable waiting state: ErrImagePull, ImagePullBackOff, InvalidImageName, ErrInvalidImage, CreateContainerConfigError, CreateContainerError, or RunContainerError | Pod stuck in terminal state — manual fix required | All |
The TLS Secret referenced by spec.certs[] is missing a required key (tls.crt or tls.key) | cert secret <namespace>/<name> is missing required key "<key>" | All |
| The CR spec fails validation during reconciliation (e.g. missing required field, invalid value) | <CR type> spec validation failed | All |
| The Queue or ObjectStorage CR referenced by an IndexerCluster or IngestorCluster cannot be found | Referenced Queue or ObjectStorage CR not found | IndexerCluster, IngestorCluster |
clusterManagerRef is empty at the point where it is required at runtime | empty Cluster Manager reference | IndexerCluster |
Detecting a terminal failure
When a terminal failure occurs, status.phase is Error and the Stalled condition flips to True. A Kubernetes Warning event with reason Stalled is also emitted and is visible in kubectl describe. Check the conditions directly:
kubectl get standalone example -o jsonpath='{.status.conditions}' | jq .
Or filter for the Stalled condition specifically:
kubectl get standalone example -o jsonpath='{.status.conditions[?(@.type=="Stalled")]}' | jq .
The Stalled condition message field describes the failure. For pod-level failures, check the pod status for more detail:
kubectl describe pod <pod-name> -n <namespace>
Recovery
Once the root cause is resolved and the operator successfully reconciles the CR, the Stalled condition is cleared and a Kubernetes Normal event with reason StalledResolved is emitted.
For a pod stuck in a terminal container state:
- Inspect the failing pod with
kubectl describe pod <pod-name> -n <namespace>to read theWaiting.ReasonandWaiting.Message. - Fix the root cause (correct the image tag, provide the missing
imagePullSecret, create the missing Secret or ConfigMap). - Delete the stuck pods — the StatefulSet controller recreates them and the operator resumes reconciliation.
kubectl delete pod <stuck-pod-name> -n <namespace>
For a malformed TLS Secret:
- Update or recreate the Secret to include both
tls.crtandtls.key. - The operator detects the fix and resumes automatically on the next reconcile cycle.
For a missing Queue or ObjectStorage CR (IndexerCluster, IngestorCluster):
- Create the missing CR in the namespace specified by the corresponding object reference, or in the cluster’s namespace when no reference namespace is set.
- The operator resumes automatically on the next reconcile cycle.
For a spec validation failure:
- Check the
Stalledconditionmessageand operator logs to identify the invalid field. - Correct the spec with
kubectl editorkubectl patch. - The operator processes the spec change and resumes reconciliation automatically.
For an empty ClusterManager reference (IndexerCluster):
- Ensure
spec.clusterManagerRef.nameis set on the IndexerCluster. - Apply the corrected spec — the operator resumes automatically.
Pause Annotations
The Splunk Operator controller reconciles every Splunk Enterprise CR. However, there might be circumstances wherein the influence of the Splunk Operator is not desired and needs to be paused. Every Splunk Enterprise CR has its own pause annotation associated with it, which when configured ensures that the Splunk Operator controller reconcile is paused for it. Below is a table listing the pause annotations:
| Customer Resource Definition | Annotation |
|---|---|
| queue.enterprise.splunk.com | “queue.enterprise.splunk.com/paused” |
| clustermaster.enterprise.splunk.com | “clustermaster.enterprise.splunk.com/paused” |
| clustermanager.enterprise.splunk.com | “clustermanager.enterprise.splunk.com/paused” |
| indexercluster.enterprise.splunk.com | “indexercluster.enterprise.splunk.com/paused” |
| ingestorcluster.enterprise.splunk.com | “ingestorcluster.enterprise.splunk.com/paused” |
| objectstorage.enterprise.splunk.com | “objectstorage.enterprise.splunk.com/paused” |
| licensemaster.enterprise.splunk.com | “licensemaster.enterprise.splunk.com/paused” |
| monitoringconsole.enterprise.splunk.com | “monitoringconsole.enterprise.splunk.com/paused” |
| searchheadcluster.enterprise.splunk.com | “searchheadcluster.enterprise.splunk.com/paused” |
| standalone.enterprise.splunk.com | “standalone.enterprise.splunk.com/paused” |
Note: Removal of the annotation resets the default behavior
Here is an example of a standalone with the pause annotation set. In this state, the Splunk Operator requeues the reconcillation without performing any reconcile operations unless the annotatation is removed.
apiVersion: enterprise.splunk.com/v4
kind: Standalone
metadata:
name: test-only-debug
namespace: splunk-operator
annotations:
standalone.enterprise.splunk.com/paused: "true"
finalizers:
- enterprise.splunk.com/delete-pvc
spec:
replicas: 1
admin-managed-pv Annotations
The admin-managed-pv annotation in the splunk-operator’s Custom Resource allows the admin to control whether Persistent Volumes (PVs) are dynamically created for the StatefulSet associated with the CR. If set to true, no PVs will be created, and the Persistent Volume Claim templates in the StatefulSet manifest will include a selector block to match app.kubernetes.io/instance and app.kubernetes.io/name labels for pre-created PVs. This means that /opt/splunk/etc and /opt/splunk/var related PVCs will contain code block like below
apiVersion: v1
kind: PersistentVolumeClaim
...
selector:
matchLabels:
app.kubernetes.io/instance: splunk-cm-cluster-manager
app.kubernetes.io/name: cluster-manager
To match selector definition like this, Persistent Volume must set labels accordingly
apiVersion: v1
kind: PersistentVolume
metadata:
name: pv-example-etc
labels:
app.kubernetes.io/instance: splunk-cm-cluster-manager
app.kubernetes.io/name: cluster-manager
When admin-managed-pv is set to false, PVs will be dynamically created as usual, providing dedicated persistent storage for the StatefulSet.
Here is an example of a Standalone with the admin-managed-pv annotation set. After
apiVersion: enterprise.splunk.com/v4
kind: Standalone
metadata:
name: single
finalizers:
- enterprise.splunk.com/delete-pvc
annotations:
enterprise.splunk.com/admin-managed-pv: "true"
PV label values
In order to prepare labels for CR’s persistent volumes you need to know values beforehand Below is a table listing app.kubernetes.io/name values mapped to CRDs | Customer Resource Definition | app.kubernetes.io/name value | | ———– | ——— | | clustermanager.enterprise.splunk.com | cluster-manager | | clustermaster.enterprise.splunk.com | cluster-master | | indexercluster.enterprise.splunk.com | indexer-cluster | | ingestorcluster.enterprise.splunk.com | ingestor-cluster | | licensemanager.enterprise.splunk.com | license-manager | | licensemaster.enterprise.splunk.com | license-master | | monitoringconsole.enterprise.splunk.com | monitoring-console | | searchheadcluster.enterprise.splunk.com | search-head | | standalone.enterprise.splunk.com | standalone |
app.kubernetes.io/instance value consist of three elements concatenated with hyphens
- “splunk”
- provided by admin CR name
- CRD kind name
For example clusterManager CR named “test” will have set app.kubernetes.io/instance as splunk-test-cluster-manager
Container Logs
The Splunk Enterprise CRDs deploy Splunkd in Kubernetes pods running docker-splunk container images. Adding a couple of environment variables to the CR spec as follows produces detailed container logs:
apiVersion: enterprise.splunk.com/v4
kind: Standalone
metadata:
name: test-only
namespace: splunk-operator
finalizers:
- enterprise.splunk.com/delete-pvc
spec:
replicas: 1
extraEnv:
- name: DEBUG
value: "true"
- name: ANSIBLE_EXTRA_FLAGS
value: "-vvvv
From the standalone above, here is a snippet from the detailed contianer log:
TASK [splunk_common : Ensure license path] *************************************
task path: /opt/ansible/roles/splunk_common/tasks/licenses/add_license.yml:15
ok: [localhost] => {
"changed": false,
"invocation": {
"module_args": {
"checksum_algorithm": "sha1",
"follow": false,
"get_attributes": true,
"get_checksum": true,
"get_md5": false,
"get_mime": true,
"path": "splunk.lic"
}
},
"stat": {
"exists": false
}
}
POD Eviction - OOM
As oppose to throttling in case of CPU cycles starvation, Kubernetes will evict a pod from the node if the pod’s memory demands exceeds the value set in the limits parameter.