Troubleshooting
Common failure modes and recovery procedures for the Virtufin API Gateway.
See also: Development guide → Troubleshooting for service-startup, port-conflict, and protoc-generation issues. Deployment guide → Troubleshooting for Kubernetes pod, Dapr sidecar, and image-pull issues.
Operation-level failures
These failures occur after the service is running, typically surfacing as gRPC status codes or HTTP 5xx responses.
UNAVAILABLE from a backend service call
Symptom: Invoke or ListMethods returns StatusCode.Unavailable for a
specific backend service.
Likely causes:
- The backend service (WorkManager, WebSocketManager) is not running, or is on
the wrong port
- The services.json config has the wrong host:port for the backend
- The Dapr sidecar is not running alongside the backend (no daprd container
in docker ps, no dapr_runtime_* log lines)
- The backend is in another namespace/pod and DNS resolution is failing
Diagnose:
# 1. Verify the backend is listening
grpcurl -plaintext <backend-host>:<grpc-port> list
# 2. Check Dapr sidecar logs
docker logs <backend-pod> -c daprd | tail -50
# 3. Check the API Gateway logs for the call stack
kubectl logs -n virtufin -l app=api | grep -i "<service-name>"
# 4. Verify the services.json config matches
cat config/services.json | jq '.services[] | select(.name=="<service-name>")'
INVALID_ARGUMENT on Subscribe
Symptom: The Subscribe RPC returns immediately with
StatusCode.InvalidArgument and the message "At least one filter must be set".
Likely cause: The request has empty services, topics, and eventTypes
arrays. The filter validation requires at least one non-empty filter.
Fix: Set at least one of services, topics, or eventTypes to a
non-empty value.
RESOURCE_EXHAUSTED on SaveState
Symptom: SaveState returns StatusCode.ResourceExhausted with a message
about state-store capacity.
Likely cause: The Dapr state store (Redis/Valkey) is full. The
default maxmemory-policy in redis.conf is noeviction, so writes fail when
memory is exhausted.
Fix: Increase the Redis instance size, or set
maxmemory-policy allkeys-lru to evict older keys.
DEADLINE_EXCEEDED on Invoke
Symptom: Invoke returns DeadlineExceeded after the configured timeout.
Likely cause: The backend method itself is slow (not the network).
Invoke has a 30-second default timeout inherited from the gRPC client. The
underlying backend may be doing expensive work (e.g., loading a large Python
worker code module).
Fix: Increase the per-RPC timeout (configurable via
Grpc:DefaultTimeoutSeconds in appsettings.json) or optimize the backend
operation.
State store failures
State changes not persisting
Symptom: SaveState returns success but the next GetState returns the old value.
Likely cause: SaveState reports failures in-band — the gRPC status is OK even
when the write failed. A client that only catches RpcException reads a failed save as a
success.
Fix: check response.status.success on every write:
curl -s http://localhost:5001/v1/state/save-state \
-H "Content-Type: application/json" \
-d '{"service":"workmanager","key":"workmanager.worker.abc","value":"{\"n\":1}"}' | jq .status
If success is false, the message is deliberately generic
("An internal error occurred") — the real cause is in the API's server logs.
Reading another service's keys succeeds unexpectedly
Symptom: GetState(service: "workmanager", key: "websocketmanager.connection.x")
returns data instead of failing.
Cause: This is expected. service selects the Dapr state store; it is not a namespace
check. Keys are not scoped to the calling service. See
security.md.
QueryState fails for every query, including {}
Symptom: All QueryState calls return Internal, even Dapr's match-everything query.
Likely cause: The configured store's Dapr component does not implement the Query API.
For Redis/Valkey this needs the RediSearch module loaded and queryIndexes metadata
declared on the component.
Fix: add the metadata and RediSearch module, or use GetBulkState against known keys.
Pub/Sub failures
Subscriptions stop receiving events
Symptom: A client subscribed via the Pubsub.Subscribe RPC stops receiving
events after a few minutes.
Likely cause: either a dropped stream or a consumer that fell behind.
SubscriptionHealthSweeperHostedService runs every minute and writes an
empty-Data heartbeat to each subscriber's channel, reaping any whose channel
cannot accept a write within 3 seconds. Each subscriber buffers up to 1024
events, so this only fires for a consumer that has genuinely stopped keeping up.
If yours does slow work per event, move it off the read loop.
Otherwise check the client side: a dropped gRPC stream (idle timeout on an
intermediary, pod restart, CancellationToken fired) ends the Subscribe call,
and the server's finally releases both the in-process subscriber and the
shared Dapr subscription. The stream does not resume itself — the client must
reconnect.
Diagnose:
# Check the sweeper logs
kubectl logs -n virtufin -l app=api | grep -i "sweeper\|heartbeat"
# Check the subscription count
curl http://localhost:5001/health
Fix: Restart the client subscription; the next sweep will recreate the health record.
Recovery procedures
Restart the gateway after a state-store outage
If the Dapr state store is unavailable for an extended period, the gateway may
have cached stale state. Restart all gateway replicas to force a fresh
ServicesConfigurationLoader.LoadAsync and rebuild the gRPC channel pool:
kubectl rollout restart deployment/api -n virtufin
kubectl rollout status deployment/api -n virtufin
Reset all in-flight subscriptions
If subscriptions are wedged (heartbeat lag, broken state), force a clean state:
# Restart all gateway replicas (closes all subscriptions)
kubectl rollout restart deployment/api -n virtufin
# Wait for old replicas to terminate
kubectl rollout status deployment/api -n virtufin
# Clients must re-subscribe
Reporting issues
If none of the above resolves your issue, gather the following and file a Gitea issue:
- API Gateway version (
LIBRARY_VERSIONfrom the build info endpoint) services.jsonconfig (redact any sensitive values)- Full error message and gRPC status code
- Relevant log lines from the gateway and the affected backend service