One faulty GPU worker failed to process requests

Partial outageResolvedVirtual Try-OnStarted Lasted 8 hours and 9 minutes

Updates

  1. Resolved

    • One GPU worker went into a bad state. It was restarted and returned to normal operation.
    • A mechanism to detect and auto-restart such states was developed and deployed.
  2. Investigating

    Some of the requests to the nightly endpoint (experimental in app) are returning errors