Skip to content

chore: evaluate Trigger.dev for scheduled cache refresh jobs #453

Description

@matheus1lva

Trigger.dev Security Analysis and Threat Model

System: Kong scheduled cache-refresh jobs
Decision: Whether to replace GitHub Actions cron execution with Trigger.dev Cloud
Prepared: 2 August 2026
Audience: Company director, engineering, and security reviewers
Status: Proposed, conditional approval

Executive summary

Trigger.dev can replace the GitHub Actions schedules used by Kong, but it does not remove third-party trust. It transfers the scheduled-job execution boundary from GitHub-hosted runners to Trigger.dev-managed workers.

The proposed jobs read production PostgreSQL data and write derived data to Redis. Trigger.dev would therefore need production database and Redis credentials, a copy of the task code and dependencies, and network access to both services. Trigger.dev would also retain operational metadata such as schedules, run status, logs, and possibly task inputs/outputs. This makes compromise of the Trigger.dev organization, deployment pipeline, worker platform, or stored secrets a path to production data access or cache corruption.

The current GitHub design already exposes equivalent credentials to an external CI platform during each run. The material security differences are:

  • Trigger.dev becomes a persistent production execution and scheduling control plane rather than an ephemeral CI runner.
  • Task code is built into a container image and deployed to Trigger.dev.
  • Secrets are stored in Trigger.dev for worker execution rather than only in GitHub Actions secrets.
  • Trigger.dev observability, task payloads, results, and checkpoints can create additional copies of data if code logs or returns sensitive values.
  • Trigger.dev dashboard users may be able to trigger, replay, cancel, deploy, or alter schedules depending on their permissions.

Recommendation: conditionally approve a limited pilot for the three existing scheduled cache jobs, provided the mandatory controls in this document are completed. Do not give Trigger.dev general-purpose production credentials, write access to authoritative Kong tables, or secrets unrelated to these jobs. Self-hosting should be considered if company policy does not permit production database credentials or operational data in a SaaS task platform.

Residual risk after the controls is assessed as Medium. Without least-privilege database credentials, MFA/access governance, safe logging, idempotency, and tested revocation, residual risk is High and production use should not be approved.

Scope

Included

The repository currently has these GitHub-scheduled production jobs:

Workflow Schedule Command Required access
refresh-cache.yml Every 30 minutes refresh-vaults.ts PostgreSQL read; Redis write
timeseries-refresh.yml Hourly timeseries/refresh.ts PostgreSQL read; Redis write
timeseries-refresh-historical.yml Daily at 00:00 UTC timeseries/refresh-historical.ts PostgreSQL read; Redis write

The manual report-refresh workflows may later use the same pattern, but are not part of initial approval.

Excluded

  • Replacing Kong's BullMQ/Redis ingestion queues or its internal 10/30-second recurring jobs
  • Moving the web application, PostgreSQL, or Redis hosting
  • CI for linting and pull requests
  • Jobs that sign blockchain transactions, control funds, or hold private keys
  • Trigger.dev self-hosting implementation

Any expansion into these areas requires a new threat review.

Proposed architecture and trust boundary

Git repository / deployment operator
              |
              | deploy bundled task code
              v
     Trigger.dev control plane
       - organization/users
       - schedules and deployments
       - encrypted environment variables
       - run metadata, logs and results
              |
              | starts isolated task worker
              v
      Trigger.dev task container
       |                    |
       | TLS                | TLS
       v                    v
Production PostgreSQL   Production Redis
 (read-only role)       (restricted cache access)
              |
              v
         Uptime Kuma
      (success/failure only)

Trigger.dev states that Cloud task code is packaged as a container image and run in an isolated environment. Its security page states SOC 2 Type II compliance, third-party penetration testing, AES-256 encryption at rest, HTTPS/TLS in transit, and optional MFA. These are useful vendor controls, but they do not replace customer-side least privilege or prevent accidental disclosure through application logs and outputs.

Assets and security objectives

Asset Objective Impact if compromised
PostgreSQL credentials and data Confidentiality; query integrity Production data disclosure, expensive queries, or unauthorized changes
Redis credentials and cache Integrity; availability Incorrect API responses, cache poisoning, deletion, or service degradation
Trigger.dev environment variables Confidentiality Direct access to the above services and monitoring endpoint
Task source and deployment artifacts Confidentiality; integrity Intellectual-property leakage or malicious production execution
Schedules and task definitions Integrity; availability Missed refreshes, duplicate work, or resource exhaustion
Logs, payloads, results, and checkpoints Confidentiality; controlled retention Persistent disclosure of data or secrets
Trigger.dev organization and API keys Authentication; authorization Unauthorized deploys, triggers, replays, and secret access
Availability and correctness of public cache data Integrity; availability Stale or incorrect customer-facing data

Threat actors

  • External attacker who compromises a Trigger.dev, GitHub, or employee account
  • Malicious or careless employee with access to the Trigger.dev organization
  • Compromised dependency, build tool, deployment token, or task container
  • Tenant-isolation failure or malicious tenant in the SaaS platform
  • Trigger.dev or subprocessors' privileged personnel
  • Automated abuse caused by a configuration error, retry storm, overlapping schedule, or faulty deployment

Threat register

Likelihood and impact are qualitative and assume production use.

ID Threat Initial risk Required mitigation Residual risk
T1 Trigger.dev account takeover permits deployments, manual runs, replays, or schedule changes High SSO if available; otherwise mandatory MFA; individual accounts; least-privilege roles; quarterly access review; immediate offboarding; alerts for deploy and schedule changes Medium
T2 Stored PostgreSQL or Redis credentials are exposed through vendor compromise, account misuse, support access, or API abuse Critical Dedicated credentials; DB read-only role; Redis ACL limited to required commands/key prefix where supported; no shared admin credentials; rotation and tested emergency revocation Medium
T3 Task code, dependency, or deployment pipeline is compromised and exfiltrates data/secrets High Lock dependencies; review task/deployment changes; protected production deploy path; dependency scanning; minimal image; egress restrictions where available; separate deploy credential Medium
T4 Sensitive values enter logs, traces, task payloads, outputs, error objects, or checkpointed memory High Pass no secrets in payloads; never log connection strings or raw rows; return only counts/status; redact errors; define retention; test log behavior with synthetic secrets Low–Medium
T5 Overlap, retries, replay, or duplicate schedules cause conflicting writes or load spikes High One concurrency slot per job; TTL for stale scheduled runs; idempotent/upsert cache writes; bounded retry with backoff; query timeouts; resource limits; no automatic retry for deterministic failures Low–Medium
T6 A compromised task modifies authoritative production data High A dedicated PostgreSQL role with SELECT only on required schemas/tables; no DDL/DML; database-side statement and connection limits Low
T7 A task poisons or deletes unrelated Redis data High Separate Redis instance/database where feasible; otherwise ACL to exact commands and key namespaces; deny FLUSH*, scripting, config, module, and broad key scans Medium
T8 Public database/Redis endpoints expand the network attack surface High Prefer Trigger.dev Private Networking via AWS PrivateLink when infrastructure is compatible; otherwise TLS, provider firewall/allowlisting, strong rotated credentials, and connection monitoring Medium
T9 Trigger.dev outage or scheduling failure leaves caches stale Medium Keep cache TTL/staleness behavior safe; Uptime Kuma dead-man monitoring; alert on missed expected run; documented GitHub/manual fallback; tested rollback Low–Medium
T10 Malicious or mistaken dashboard action changes a production schedule Medium Define schedules declaratively in reviewed code; prohibit imperative production schedules except emergency use; audit changes; restrict production project membership Low–Medium
T11 Cross-tenant isolation failure exposes code, secrets, task state, or network routes High Review SOC 2 report and penetration-test summary under NDA; vendor due diligence; PrivateLink tenant-isolation validation; minimize data sent to the service; contractual incident terms Medium
T12 Vendor/subprocessor or data-location terms conflict with company requirements High Legal review of DPA, subprocessors, transfer mechanism, retention/deletion, and breach terms; document approved data classification Medium
T13 Secret remains usable after migration, employee departure, or incident High Inventory owner and expiry for every credential; rotate at migration; revoke GitHub copies after rollback window; quarterly rotation; incident runbook tested before launch Low–Medium
T14 Observability or monitoring URL leaks through logs and permits forged status updates Medium Treat push URLs as secrets; only send fixed status/message fields; rotate on disclosure; do not include production data Low
T15 Cost/resource denial through repeated triggers or runaway jobs Medium Concurrency limits, max duration, query timeout, billing alerts, run-rate alerts, and manual kill/revoke procedure Low–Medium

Mandatory controls before production

Approval is conditional on all of the following:

  1. Dedicated identities: Create a new PostgreSQL login for Trigger.dev with SELECT only on the tables/views actually queried. It must have no write, DDL, ownership, role-management, replication, or bypass-RLS privileges. Create a dedicated Redis identity with the narrowest feasible key and command ACL.
  2. Transport and network: Require verified TLS for PostgreSQL and Redis. Use AWS PrivateLink if the relevant infrastructure is in AWS and the plan supports it. If public endpoints are unavoidable, document firewall controls and confirm whether Trigger.dev offers stable egress IPs for the selected region/plan before launch.
  3. Separate production project/environment: Production tasks, credentials, and schedules must be isolated from development and previews. Development code must not inherit production environment variables.
  4. Strong access control: Require MFA for every Trigger.dev user. Prefer SSO and role-based access on a plan that supports the required controls. No shared accounts. Restrict production deployment and environment-variable management to named maintainers.
  5. Reviewed, declarative schedules: Store production cron definitions in code and deploy them through review. Do not rely on mutable dashboard-only schedules for normal operation.
  6. Safe task semantics: Configure per-job concurrency of one, stale-run TTL, bounded retries, max duration, database statement timeout, and idempotent cache replacement. Verify that replay cannot corrupt state.
  7. Data minimization: Scheduled tasks must have no business-data payload. Logs and results must contain only job name, timing, counts, and sanitized errors. Never return rows, cache contents, credentials, or connection strings.
  8. Credential handling: Mark all credential variables as secret. Do not bulk-upload the repository .env. Upload only an allowlist of required variables. Rotate new credentials after the pilot setup and again on any suspected disclosure.
  9. Monitoring: Alert on failure, excessive duration, repeated retry, unexpected manual run, schedule/deployment change, and missed run. Preserve Uptime Kuma checks or implement an equivalent independent dead-man signal.
  10. Fallback and revocation: Document how to disable Trigger.dev schedules, revoke its DB/Redis users, restore GitHub schedules, and verify cache freshness. Exercise this procedure in staging.
  11. Vendor review: Security/legal must review the current SOC 2 Type II report, recent penetration-test summary, DPA, subprocessor list, support-access model, incident notification terms, backup/retention behavior, deletion process, and role/audit-log availability for the selected plan.
  12. No signing authority: Do not place wallet private keys, treasury credentials, deployer keys, or any credential capable of moving funds in Trigger.dev under this approval.

Recommended implementation pattern

  • Create three small Trigger.dev scheduled tasks corresponding one-to-one with the existing GitHub scheduled workflows.
  • Import only the cache-refresh implementation needed by each task; do not package unrelated operational scripts or local .env files.
  • Use declarative UTC schedules to preserve the current behavior.
  • Set concurrency to one per cache family. A delayed run should expire rather than accumulate behind an already-running refresh.
  • Use a restricted, read-only PostgreSQL role and an isolated cache-writer Redis identity.
  • Return a fixed result such as { ok: true, recordsProcessed: number }; sanitize all thrown errors.
  • Deploy to staging first using non-production credentials and representative data.
  • Run Trigger.dev and GitHub in observation mode without both writing the same production keys. Compare outputs, duration, failure behavior, and monitoring for at least one week.
  • Cut over one low-impact job first. Keep the GitHub workflow manually dispatchable but remove its schedule. After a defined rollback window, remove production secrets from GitHub if they are no longer required there.

Vendor assurances and open due-diligence items

Public Trigger.dev material states:

  • SOC 2 Type II compliance; the report is available on request to Enterprise customers.
  • Regular independent penetration testing; the report is available on request to Enterprise customers.
  • AES-256 encryption at rest and HTTPS/TLS encryption in transit.
  • MFA availability.
  • A DPA under which Trigger.dev acts as processor and commits to breach notification within 72 hours of awareness for personal-data breaches.
  • Multi-tenant storage in AWS us-east-1, with logical customer separation at the application layer, according to the DPA.
  • Private AWS connectivity via PrivateLink on Pro and Enterprise plans, with per-organization network policy and AWS-level authorization described by the vendor.
  • A broad and changeable subprocessor list. Legal and security should confirm which subprocessors actually receive task code, secrets, payloads, logs, or operational metadata.

The public pages do not, by themselves, resolve every control needed for approval. Obtain written answers to:

  1. What are the retention periods for run payloads, outputs, logs, container images, failed runs, backups, and CRIU checkpoints?
  2. Are environment-variable values included in backups, and how quickly are deleted/rotated values removed from replicas and backups?
  3. Which Trigger.dev personnel can access customer secrets, task containers, logs, and checkpoints, under what approval process, and is customer-visible audit history available?
  4. What production roles and audit events are available on the intended plan?
  5. Are stable outbound IP addresses available if PrivateLink cannot be used?
  6. Can task egress be limited to the database, Redis, and Uptime Kuma endpoints?
  7. What regions are available for task execution and control-plane data, and can residency be contractually fixed?
  8. Which subprocessors process the task execution data for this specific service configuration?
  9. What recovery objectives and historical uptime commitments are included in the selected plan?
  10. Can the company receive security advisories and advance notice of subprocessor changes?

Comparison with the current GitHub Actions design

Area GitHub scheduled workflow Trigger.dev Cloud Assessment
Execution Ephemeral hosted CI runner Managed task worker/container Similar third-party code execution; different vendor boundary
Secret storage GitHub Actions secrets Trigger.dev secret environment variables Additional secret store unless GitHub copies are removed
Scheduling YAML cron in repository Declarative code or mutable dashboard schedule Prefer declarative code to retain reviewability
Observability Action logs and artifacts Run logs, traces, results, dashboard, possibly checkpoints Trigger.dev may retain more task state; minimize output and define retention
Retries/replay Workflow-specific/manual First-class retries and replay Operational improvement, but duplicate-write risk rises
Access GitHub repository/org roles Separate Trigger.dev organization roles New identity lifecycle and attack surface
Network GitHub runner to public service endpoints Trigger.dev worker; PrivateLink available on some plans Potential improvement if PrivateLink is used
Availability dependency GitHub Actions Trigger.dev scheduler/control plane Vendor dependency changes rather than disappears

Acceptance criteria and decision gate

Production cutover may proceed only when:

  • Every mandatory control has an assigned owner and evidence.
  • Security/legal vendor due diligence is accepted for the data classification involved.
  • Staging proves least-privilege access, sanitized logs, bounded retries, non-overlap, monitoring, and revocation.
  • A production pilot completes without correctness or availability regression.
  • The system owner accepts the documented Medium residual risk.

The decision must be revisited if Trigger.dev will process personal/confidential payloads, gain write access to authoritative databases, execute ingestion queues, receive signing keys, or become unavailable beyond the agreed recovery window.

References

Review record

Role Name Decision/date
Engineering owner
Security reviewer
Legal/privacy reviewer
Director/risk owner

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions