Logging & Events
Logging
The operator uses Go’s log/slog package for structured logging. The global logger is configured once at startup in cmd/main.go via logging.SetupLogger() and set as the default with slog.SetDefault().
Getting a Logger
| Where | How |
|---|---|
Controller Reconcile | logger := slog.Default().With("controller", "Name", "name", req.Name, "namespace", req.Namespace, "reconcileID", controller.ReconcileIDFromContext(ctx)) |
| Business logic (pkg) | logger := logging.FromContext(ctx).With("func", "FunctionName") |
Controllers inject the logger into context so downstream code can retrieve it:
logger := slog.Default().With("controller", "Standalone", "name", req.Name, "namespace", req.Namespace, "reconcileID", controller.ReconcileIDFromContext(ctx))
ctx = logging.WithLogger(ctx, logger)
The reconcileID is a unique identifier assigned by controller-runtime to each reconcile invocation. It is automatically propagated to all downstream log calls via context, making it possible to correlate every log line produced during a single reconcile pass — even across deeply nested function calls.
Log Levels
| Level | Use for |
|---|---|
Debug | Verbose diagnostics (state dumps, intermediate values) |
Info | Normal operations (reconcile start, requeue, status changes) |
Warn | Recoverable issues (deprecated config, fallback behavior) |
Error | Failures that need attention (API errors, broken invariants) |
Configure at runtime via LOG_LEVEL env var or --log-level flag. Values: debug, info, warn, error.
Writing Log Messages
Message format: short, lowercase, past tense or present participle. Describe what happened, not the function name. Use PascalCase for CRD type names in messages — ClusterManager, IndexerCluster, SearchHeadCluster, LicenseManager, LicenseMaster, ClusterMaster, MonitoringConsole, Standalone.
// Good
logger.InfoContext(ctx, "ClusterManager types not found in namespace", "error", err)
logger.InfoContext(ctx, "scaling up IndexerCluster", "replicas", replicas)
// Bad - inconsistent casing
logger.InfoContext(ctx, "clusterManager types not found")
logger.InfoContext(ctx, "scaling up indexer cluster")
// Good
logger.InfoContext(ctx, "statefulset updated", "replicas", replicas)
logger.ErrorContext(ctx, "failed to fetch secret", "error", err, "secret", name)
// Bad - don't restate the function name or use generic messages
logger.InfoContext(ctx, "ApplyStandalone")
logger.InfoContext(ctx, "error occurred")
Rules:
- Always use
*Context(ctx, ...)variants (InfoContext,ErrorContext, etc.). - Always pass
"error", erras a key-value pair — never as the message string. - Use consistent key names across the codebase. Key names must be
camelCase— no spaces, no Title Case, no special characters like parentheses.
// Good
logger.InfoContext(ctx, "start", "crVersion", instance.GetResourceVersion())
logger.InfoContext(ctx, "Requeued", "periodSeconds", int(result.RequeueAfter/time.Second))
logger.InfoContext(ctx, "Getting Apps list", "bucket", client.BucketName)
// Bad - spaces and inconsistent casing in keys
logger.InfoContext(ctx, "start", "CR version", instance.GetResourceVersion())
logger.InfoContext(ctx, "Requeued", "period(seconds)", int(result.RequeueAfter/time.Second))
logger.InfoContext(ctx, "Getting Apps list", "AWS S3 Bucket", client.BucketName)
Common key names:
| Key | Meaning |
|---|---|
"error" | The error value |
"name" | CR or resource name |
"namespace" | Kubernetes namespace |
"controller" | Controller name |
"func" | Function name (in pkg layer) |
"reconcileID" | Unique ID per reconcile pass (set by controller) |
"replicas" | Replica count |
"phase" | CR phase |
"podName" | Pod name |
"appName" | App name |
"bucket" | Storage bucket name (S3/GCS/Azure) |
"crVersion" | CR resource version |
"periodSeconds" | Requeue delay in seconds |
- Don’t log and return the same error. controller-runtime automatically logs every non-nil error returned from
Reconcile()vialog.Error(err, "Reconciler error"). If you also log it in the business logic, the same error appears twice. Instead, wrap the error with context usingfmt.Errorf("description: %w", err)so the single controller-runtime log line is descriptive.
Error Wrapping
The Apply* functions in pkg/splunk/enterprise/ wrap validation errors before returning them. This adds context about which resource type failed, producing a clear chain when controller-runtime logs the final error:
validate clustermanager spec: license not accepted, please adjust SPLUNK_GENERAL_TERMS ...
There are two distinct error-return patterns depending on whether the failure is terminal (user-actionable, no requeue) or transient (retryable):
Terminal errors (spec validation, missing required CRs, malformed secrets) — use splcommon.NewTerminalError:
err = validateClusterManagerSpec(ctx, client, cr)
if err != nil {
eventPublisher.Warning(ctx, EventReasonValidateSpecFailed,
fmt.Sprintf("Spec validation failed for %s — check operator logs", cr.GetName()))
setPhaseAndConditions(enterpriseApi.PhaseError, "Cluster Manager spec validation failed")
return reconcile.Result{}, splcommon.NewTerminalError(EventReasonValidateSpecFailed, "Cluster Manager spec validation failed", err)
}
Transient errors (API blips, config apply failures) — use fmt.Errorf with %w:
namespaceScopedSecret, err := ApplySplunkConfig(ctx, client, cr, cr.Spec.CommonSplunkSpec, SplunkIndexer)
if err != nil {
eventPublisher.Warning(ctx, EventReasonApplySplunkConfigFailed,
fmt.Sprintf("Failed to apply general config for %s — check operator logs", cr.GetName()))
setPhaseAndConditions(enterpriseApi.PhaseError, "Failed to apply configuration")
return result, fmt.Errorf("apply splunk config: %w", err)
}
Rules:
- Use
splcommon.NewTerminalError(reason, message, cause)for failures that require user intervention and must not be retried automatically. The controller layer intercepts terminal errors and setsStalled=Trueon the CR status automatically. - Use
fmt.Errorf("description: %w", err)with%w(not%v) for transient failures so callers can inspect the error chain witherrors.Is/errors.Unwrap. The prefix should be lowercase and describe the operation, e.g."apply splunk config","update monitoring console configmap". - Don’t wrap and log the same error — pick one. Wrapping is preferred because it keeps context in the error chain and avoids double-logging.
// Good — wrap and return, no explicit log needed
return result, fmt.Errorf("apply splunk config: %w", err)
// Bad — double-logged: once here, once by controller-runtime
scopedLog.Error(err, "Failed to apply splunk config")
return result, err
When writing tests that assert on wrapped errors, use strings.Contains rather than an exact match so the test doesn’t break if the wrapping prefix changes:
// Good — resilient to wrapping changes
if !strings.Contains(err.Error(), "license not accepted") {
t.Errorf("Unexpected error: %v", err)
}
// Bad — brittle, breaks if wrapping prefix is renamed
if err.Error() != "validate standalone spec: license not accepted, ..." {
t.Errorf("Unexpected error: %v", err)
}
- Sensitive data (passwords, tokens, secrets) is automatically redacted by the handler.
Configuration
| Env Var | Flag | Default | Description |
|---|---|---|---|
LOG_LEVEL | --log-level | info | Minimum log level |
LOG_FORMAT | --log-format | json | json or text |
LOG_ADD_SOURCE | --log-add-source | false | Include source file/line (auto-enabled at debug) |
Flags take precedence over env vars.
Kubernetes Events
Events are user-facing signals visible via kubectl describe. Use them for significant state changes that an operator user should see — not for internal debugging.
How It Works
- Controllers pass the event recorder into context:
ctx = context.WithValue(ctx, splcommon.EventRecorderKey, r.Recorder) - Business logic retrieves a publisher:
eventPublisher := GetEventPublisher(ctx, cr) - Emit events using constants from
pkg/splunk/enterprise/event_reasons.go:eventPublisher.Normal(ctx, EventReasonScaledUp, fmt.Sprintf("Successfully scaled %s from %d to %d replicas", cr.GetName(), old, new)) eventPublisher.Warning(ctx, EventReasonValidateSpecFailed, fmt.Sprintf("Spec validation failed for %s — check operator logs", cr.GetName()))
Event Reason Constants
All event reasons are defined as constants in pkg/splunk/enterprise/event_reasons.go. Never use string literals for event reasons — always use the EventReason* constants. This ensures consistency across controllers and makes reasons searchable.
Normal reasons:
| Constant | Value | Use for |
|---|---|---|
EventReasonScaledUp | ScaledUp | Successful scale up |
EventReasonScaledDown | ScaledDown | Successful scale down |
EventReasonClusterInitialized | ClusterInitialized | Cluster first becomes ready |
EventReasonClusterQuorumRestored | ClusterQuorumRestored | Quorum recovered |
EventReasonPasswordSyncCompleted | PasswordSyncCompleted | Secret sync finished |
EventReasonStalledResolved | StalledResolved | Stalled condition cleared — reconciliation has resumed |
Warning reasons (common):
| Constant | Value | Use for |
|---|---|---|
EventReasonValidateSpecFailed | ValidateSpecFailed | CR spec validation errors |
EventReasonApplySplunkConfigFailed | ApplySplunkConfigFailed | General config apply failures |
EventReasonAppFrameworkInitFailed | AppFrameworkInitFailed | App framework init errors |
EventReasonApplyServiceFailed | ApplyServiceFailed | Service create/update failures |
EventReasonStatefulSetFailed | StatefulSetFailed | StatefulSet get failures |
EventReasonStatefulSetUpdateFailed | StatefulSetUpdateFailed | StatefulSet update failures |
EventReasonDeleteFailed | DeleteFailed | CR deletion failures |
EventReasonSecretMissing | SecretMissing | Required secret not found |
EventReasonCertSecretMalformed | CertSecretMalformed | TLS Secret missing required key |
EventReasonResolveQueueObjectStorageFailed | ResolveQueueObjectStorageFailed | Terminal error reason when the referenced Queue or ObjectStorage CR is not found — sets Stalled=True with message "referenced Queue or ObjectStorage CR not found", no requeue. Other failures from the same path (transient API errors, ConfigMap/Secret write failures) are retryable. The accompanying Warning event uses the literal reason EnsureDefaultsFailed |
EventReasonEmptyClusterManagerRef | EmptyClusterManagerRef | ClusterManagerRef is empty during reconciliation |
EventReasonUpgradeCheckFailed | UpgradeCheckFailed | Upgrade path validation errors |
EventReasonStalled | Stalled | Stalled condition onset — manual intervention required |
See event_reasons.go for the full list.
To add a new event reason: add a constant to event_reasons.go, then use it in the code.
When to Use Events vs Logs
| Scenario | Log | Event |
|---|---|---|
| Reconcile start/requeue | Yes | No |
| API call failed (transient) | Yes | No |
| Spec validation failed | Yes | Yes (Warning) |
| Successful scale up/down | Yes | Yes (Normal) |
| Phase transition | Yes | Yes (Normal) |
| Security-related failure | Yes | Yes (Warning) |
Writing Event Messages
Reason: always use an EventReason* constant — e.g. EventReasonScaledUp, EventReasonValidateSpecFailed.
Message: sentence case, user-friendly, include the resource name. Never include raw err.Error() in events — error details belong in logs only. Events are visible to any user with kubectl describe access and may leak internal paths, secret names, or stack traces. Instead, write a user-actionable summary and point to logs for details.
// Good — uses constant, event gives a user-actionable summary, log has full error
logger.ErrorContext(ctx, "smartstore volume key validation failed", "error", err)
eventPublisher.Warning(ctx, EventReasonRemoteVolumeKeyCheckFailed,
fmt.Sprintf("Remote volume key change check failed for %s — check operator logs", cr.GetName()))
eventPublisher.Normal(ctx, EventReasonScaledUp,
fmt.Sprintf("Successfully scaled %s from %d to %d replicas", cr.GetName(), oldCount, newCount))
// Bad — string literal instead of constant
eventPublisher.Warning(ctx, "ValidationFailed", "...")
// Bad — leaking err.Error() into events
eventPublisher.Warning(ctx, EventReasonValidateSpecFailed,
fmt.Sprintf("Invalid smartstore config: %s", err.Error()))
// Bad — too vague, no context
eventPublisher.Warning(ctx, "Error", "something went wrong")
Log + Event combo for errors:
err = validateStandaloneSpec(ctx, client, cr)
if err != nil {
// Log: full error for operator developers / debugging
logger.ErrorContext(ctx, "spec validation failed", "error", err)
// Event: user-actionable summary visible in kubectl describe
eventPublisher.Warning(ctx, EventReasonValidateSpecFailed,
fmt.Sprintf("Spec validation failed for %s — check operator logs", cr.GetName()))
return result, err
}
Rules:
- Always use
EventReason*constants — never string literals for event reasons. - Use
Normalfor successful operations the user should know about. - Use
Warningfor failures or degraded states that may need user intervention. - Always include the resource name in the message.
- Never pass
err.Error()into event messages — log the full error, keep events clean. - Don’t emit events for routine internal operations — keep the event stream actionable.
Terminal Errors
Some failure conditions cannot self-heal without external intervention — for example, a pod stuck in ImagePullBackOff, a TLS Secret with a missing key, or a CR spec that fails validation. The operator wraps these in splcommon.NewTerminalError(reason, message, err) before returning from the Apply* function. controller-runtime treats terminal errors as non-retriable: the CR is not requeued and the reconcile loop stops.
Terminal failures are surfaced to the user via the Stalled status condition (Stalled=True) in addition to phase=Error. See Status Conditions for the full condition schema.
How terminal errors propagate
Business logic signals a non-retriable failure by returning splcommon.NewTerminalError(reason, message, cause) directly — no event publisher helpers are needed. The controller layer intercepts any terminal error returned from Reconcile() and automatically calls splcommon.UpsertStalledCondition to set Stalled=True on the CR status.
Four-layer propagation pattern:
// Layer 1 — business logic (e.g. certs.ValidateCertSecret):
// returns splcommon.NewTerminalError(reason, message, cause) for user-actionable failures.
// Layer 2 — getXxxStatefulSet:
// fmt.Errorf with %w preserves the TerminalError in the chain — no explicit check needed.
certMounts, err := certs.ReconcileCerts(ctx, client, cr, toCertEntries(cr.Spec.Certs))
if err != nil {
return nil, fmt.Errorf("reconcile certs: %w", err)
}
// Layer 3 — Apply* function:
// emits event and sets phase, then returns result + err; terminal vs transient is the controller's concern.
statefulSet, err := getStandaloneStatefulSet(ctx, client, cr)
if err != nil {
eventPublisher.Warning(ctx, EventReasonStatefulSetFailed, "get standalone statefulset failed — check operator logs")
setPhaseAndConditions(enterpriseApi.PhaseError, "Failed to create or update StatefulSet")
return result, err
}
// Layer 4 — controller Reconcile:
// returns reconcile.Result{} (no requeue) for terminal errors, result for transient ones.
if _, ok := splcommon.TerminalMessage(err); ok {
return reconcile.Result{}, err
}
return result, err
Layer responsibilities:
- Business logic (
certs.ValidateCertSecret, etc.) — returnssplcommon.NewTerminalErrorfor non-retryable failures. The error’sReasonandMessageare structured for theStalledcondition; full detail lives inErrfor operator logs. getXxxStatefulSet— wraps errors withfmt.Errorf("...: %w", err). No explicit terminal check needed:%wpreserves the full error chain, soTerminalErrorremains detectable viaerrors.As/splcommon.TerminalMessageat any point up the stack.Apply*— emits aWarningevent, setsphase=Error, returnsresult, err. Does not inspect for terminal errors — that is the controller’s responsibility.- Controller
Reconcile— callssplcommon.TerminalMessage(err)to upsert theStalledcondition, then returnsreconcile.Result{}, errfor terminal errors (stops requeueing) orresult, errfor transient ones.
Callers must not inspect specific sub-package error types (e.g. certs.ErrCertSecretMalformed) — the sub-package is responsible for returning a terminal error when appropriate.
Rules:
- Return
splcommon.NewTerminalError(reason, message, cause)from business logic — the controller setsStalled=Trueautomatically. - Never return
reconcile.TerminalErrorfor transient failures (network blips, temporary API unavailability). - Use
splcommon.IsStalled(cr.Status.Conditions)to check programmatically whether a CR is currently stalled.
Stalled condition events
After computing the new Stalled condition (and before persisting it), every enterprise controller calls enterprise.EmitStalledTransitionEvents. A Warning event fires on every reconcile where Stalled=True; the Normal recovery event fires only when the condition transitions from True→False:
| Condition | Event type | Reason constant |
|---|---|---|
Stalled=True (every terminal reconcile) | Warning | EventReasonStalled |
Stalled=True → Stalled=False | Normal | EventReasonStalledResolved |
These events are visible via kubectl describe <cr-type> <name> and can be used in alerting pipelines to detect stalls without polling the condition array. EmitStalledTransitionEvents is defined in pkg/splunk/enterprise/events.go.
Terminal container waiting states detected by splctrl.CheckPodsForTerminalFailures:
Container Waiting.Reason | Cause |
|---|---|
ErrImagePull | Image pull failed (bad credentials or unreachable registry) |
ImagePullBackOff | kubelet backing off after repeated pull failures |
InvalidImageName | Image reference is syntactically malformed |
ErrInvalidImage | Image reference resolves but is not a valid image |
CreateContainerConfigError | env-var or volume references a missing ConfigMap or Secret key |
CreateContainerError | OCI runtime cannot create the container |
RunContainerError | OCI runtime cannot run the container (invalid entrypoint, missing binary) |