API key rotation with zero downtime

From a governance perspective, we are required to rotate API credentials every 90 days per our information security policy. It is worth noting that we currently have seven production integrations consuming the FieldPulse API, including our customer portal, billing reconciliation service, and two vendor data feeds.

I have confirmed that the API key management interface supports only a single active key per account. When generating a new key, the previous key is immediately invalidated. This presents a significant operational risk: any credential rotation will result in service interruption until all seven integration points are updated and redeployed.

I am seeking clarification on the following points:

  1. Does FieldPulse support multiple concurrent API keys, or is there a mechanism to stage a new key before invalidating the existing one?
  2. If not, what is the recommended approach for credential rotation in high-availability environments?
  3. Are there plans to implement scoped API keys with independent lifecycle management?

Our current workaround involves scheduling rotation during a maintenance window, but this does not satisfy our continuous availability requirements. I would appreciate any guidance on how other organizations have addressed this constraint.

Parents
  • Hey Fatima — yeah, this one is a real pain point and I'm not going to sugarcoat it. The short answer is: no zero-downtime rotation today with a single account key. The key swap is atomic and the old one dies immediately.

    That said, here are the paths I've seen teams take:

    Option 1: Dual Account Strategy (Most Common)

    Spin up a second FieldPulse account as a "staging" environment. Generate your new key there, pre-configure your integrations to read from a key management service (KMS) or secrets manager, and do a hot swap at the application layer. Not elegant, but it works.

    Option 2: Application-Level Key Caching

    Some teams implement a brief overlap by having their services accept both keys temporarily — old key in memory, new key ready to pick up on next deploy. This requires you to control the validation logic, which obviously only works for integrations you own end-to-end.

    Option 3: The Maintenance Window (I know, I know)

    Scoped keys with independent rotation are on our roadmap for H2 — I can say that with some confidence because I'm literally looking at the Jira epic. No firm ETA, but it's getting real design attention.

    For now, I'd recommend documenting your rotation procedure in a runbook and accepting the 5–10 minute blip. Not great, but mitigatable with retry logic and queueing on your side.

    Want me to dig into any of these approaches? I can share some Terraform snippets for the dual-account setup if that's helpful.

Reply
  • Hey Fatima — yeah, this one is a real pain point and I'm not going to sugarcoat it. The short answer is: no zero-downtime rotation today with a single account key. The key swap is atomic and the old one dies immediately.

    That said, here are the paths I've seen teams take:

    Option 1: Dual Account Strategy (Most Common)

    Spin up a second FieldPulse account as a "staging" environment. Generate your new key there, pre-configure your integrations to read from a key management service (KMS) or secrets manager, and do a hot swap at the application layer. Not elegant, but it works.

    Option 2: Application-Level Key Caching

    Some teams implement a brief overlap by having their services accept both keys temporarily — old key in memory, new key ready to pick up on next deploy. This requires you to control the validation logic, which obviously only works for integrations you own end-to-end.

    Option 3: The Maintenance Window (I know, I know)

    Scoped keys with independent rotation are on our roadmap for H2 — I can say that with some confidence because I'm literally looking at the Jira epic. No firm ETA, but it's getting real design attention.

    For now, I'd recommend documenting your rotation procedure in a runbook and accepting the 5–10 minute blip. Not great, but mitigatable with retry logic and queueing on your side.

    Want me to dig into any of these approaches? I can share some Terraform snippets for the dual-account setup if that's helpful.

Children
No Data