Scheduled jobs failing (again) (again 😄) (Ongoing Known Issue)

Is this the acknowledgment of this issue?

Investigating - Some North American users may be experiencing issues with loading resources in the mobile app and web UI, arming/disarming Smart Home Monitor, and the execution of SmartApps. The engineering team is working on the issue and we will provide an update shortly.
Oct 14, 14:19 EDT

If so, can someone articulate what is causing this?

Really annoying… this morning…

  1. Mode and/or SHM status did not change, when CoRE Piston executed. Re-Ran piston several times and did not fix it.
  2. Hub displayed a notification the hub went offline, which if it did it was the ST Hub or the Cloud, not my internet. Amazingly, the app showed the hub was online (live status) and turning on a light with the app worked.
  3. Lights took a very very very long time to turn on in response to motion (smart lighting) so I manually rebooted the hub.

Sadly, this ended a good 75 day run I had with basically no issues.

Establishes again that ST can’t be relied upon for important functions. A huge reminder, it’s entirely possible the system will not inform me when an intrusion occurs even if every other required support system is up and running (power, internet, etc) - and perhaps even worse could set off a false intrusion and siren, etc and disturb or scare the family.

ELI5?
:blush:

Same @JH1 I seemingly have been very stable and when people were complaining my setup had been very reliable since finding the Cree/Osram bulb bug. Switching most my bulbs to the Hue Bridge had made my home like dang near 100% for months now, easily 3 months. Last few days things have been super slow, aka hue delay. And then starting last night tons of failures. Routines not firing, stock basic automations not working or partially working. yet the status page is ā€œinvestigatingā€. I love the spin on it, dont say we have a problem… Hell the IDE has been unusable for me… Cant view Hub, events, and other parts of the page just give 500’s. But hey, lets investigate.

While I am both curious and of course annoyed by additional outages, to put it in perspective, if all I endure is what happened this morning I will survive. Relatively speaking, comparing to some of the outages and degradation I have been through.

Investigating is good, acknowledging there is an issue and seeking to understand cause is good. Let’s see if discovered cause is communicated clearly and if the issue remains on status until truly resolved…

Litmus test…

@bamarayne @bridaus

Report Status!

About a week ago my mode didn’t change once… and I believe it was ST. That’s the only problem I’ve had in weeks (months maybe?). I have a thermostat that was acting wonky but I’m pretty sure it wasn’t ST because my other two were fine.

@bridaus
Try to go to IDE and click on hubs… I get:
Oh No! Something Went Wrong!
Error
500: Internal Server Error

I also have issues going to a Thing in the mobile app and viewing the device logs in the Recently tab… takes forever, fails,etc…

Check it out, let us know results…

Update - While the connectivity and SmartApp execution issues have improved, we are continuing to investigate additional performance improvements.
Oct 14, 22:30 EDT

Okay, can someone share the information that backs up this statement?

Still seems like a mess to me now I am losing state of my pistons. Having to rebuild them all as my "if"s are magically empty.

I am beginning to think the status update can be translated to:

ā€˜Curious issue. We’re tired, going to bed. Catch up Monday’

Yea - one of the issues was ā€œfixedā€ by replacing 3 nodes in our events cluster for na01. These are read timeouts from that cluster. Reducing this should have helped with the UI, IDE, smartapp executions, etc…

The other issue that we’re seeing is timeouts for saving events:

It doesn’t look like the replacement fixed the issue here as the timeouts are still elevated and follow a cyclical pattern. (Can see other Cassandra metrics trending upwards again). This mainly affects execution and the chance of it occurring increases in proportion to the rate of created events per app execution. So while many smartapps/devices are firing fine certain ones are hit more often (the ones that create more events). We have a change that can go out to make the event creation async which would help with executions but that could have a number of unintended side affects and would change the behavior of the core part of the system - it would be much safer to figure out what happened in the last couple of days to cause these spikes.

I haven’t seen a lot of problems.

The app had been slow. I’ve gotten a few timeouts unable to save pages. Other than that the app had been ok.

The ide had been just fine. I’ve been in it literally all day and night working on a project. No issues there at all.

My system routine to change the mode to night mode at 2030 tonight failed to change the mode. My mode manager piston detected it and corrected the mode within a couple of minutes. I only knew it haired because the mode manager sends me a message.

Other than that my system had been running spot on.

My widgets that run some of my routines have disappeared from my iPhone as well. I had to set them up again.

This is the issue I’ve seen the last few days. So far, automations are fine, things are just sluggish updating on the mobile app.

Still happening this morning…

It’s probably your project that is causing the issue, see Vlad’s charts.

Any correlation to Project Frankenstein?

Thanks @vlad

Once again, appreciate the straight forward responses.

Are these metrics and charts monitored in semi-real time or were these put together as way to tease metrics out as part of the troubleshooting effort? Obviously there a countless metrics that could be captured and monitored, I am just wondering if these or others are being monitored proactively because it seems like this caught ST by surprise when it was reported by the community but the metrics demonstrate there was clearly something alarming going on long before then…

@Barkis
Completely antidotal, but my good morning routine seemed to execute completely and very timely this morning… ?
My hubs in IDE loaded up.
Recently tab on several devices loaded reasonably fast.

Agree: switching between hubs seems faster, although switching between ā€˜activity feed’ and ā€˜messages’ under hub ā€˜notifications’ is still sluggish compared to what I remember. YMMV, of course…

Had error 500 in IDE yesterday, support said it was the known issue. Last three days and now this morning things are randomly slow, seven seconds to turn on my bathroom light this morning…

Had a few missed things as well, so far stuff other than simple light activity is working right now.

Those specific ones I put together after looking at a few failed executions and seeing how the database issues were impacting the platform - they are counts of specific exceptions. There were other metrics (and in hindsight obvious ones) that should have spurred the investigation pretty much right after the issues started though they weren’t ones we alert one. As @JDRoberts said though - a failure to a user is a failure, doesn’t matter what the cause was and there is something seriously wrong with our processes if the community is flagging issues like this…

so when should it be fixed is the question? My home automations have been super slow this week .

Can’t get to the hub tab of the IDE to reset one of my borked SA’s, that borked itself today it would seem. Now I can’t control any of my Insteon devices. That’s a bit of a problem…

Edit: Error 500 still.