They are on AWS and most likely using KVM as the hypervisor.
Ditching SHM is not an option if you are paying for Scout Monitoring, as far as I know.
Personally, I’ve dealt with infrastructure bugs in environments that are complicated, and I can accept that these things are tricky. If a runaway process is eating all the database resources, adding resources won’t fix it.
My question is what engineering practices will be taken on to mitigate these problems in the future? There will always be bugs. Will the testing environment be augmented? Will unit testing be enforced? Will integration testing be automated? Will feature switches be used to control rollouts? Will rollouts happen in stages, so that changes can be analyzed on a small subset of users before making their way to a larger population?
If there was a multi-day outage (and it’s exact length is in question) something should be changed to prevent a recurrence, and the changes should also be designed to enhance the overall reliability and stability of the system.
I hope that’s true but doesn’t seem like that’s the case with his statement. Not a kickstarter supporter but been with ST long enough to know all the instabilities and starting losing faith with ST.
@slagle @bakken7 , I hate to be the constant negitive voice but this is the 3rd or 4th time we have heard this excuse from the ‘Engineering’ team. Each time you guys blame the DB, need to migrate to a new DB engine, have to grow the cluster, etc. Wasn’t it just a month or so ago you migrated to a new DB cluster and lost most of our historical data?
I’m extra critical against any cloud provider as this is my area of expertise and I can easily see though the poor planning / execution / inexpirence (I’ve been doing this for a long time for some really big companies).
I made the suggestion before and will say it over and over HIRE TALENT THAT KNOWS WHAT THEY ARE DOING. Where is the capacity planning? Where are the performance metrics? Where’s the change control process? Where’s the rollback / contingency plan? None of these things exist at SmartThings because you are fully focused around the DevOps/Development life cycle and haven’t taken in account what is needed to be a true SaaS/Cloud provider. Anyone can write code - you need talented, EXPERIENCED engineers from their relevant fields, not just developers.
There is no way in this word this issue should have taken a week+ to identify. A good DBA can tell you the SECOND the database isn’t performing. They can tell you if it’s network, CPU, memory or storage that is at fault - almost instantly! A good server admin can predict resource utilization, trend usage and make capacity plans. A good network engineer can monitor traffic flows, analyze patterns and identify heavy usage. Or I know at least me and my teams can.
</rant> (oh look, I can code too)
We’ve been told “yes” to all the above questions for over 2 years (probably most demanded by Customers and emphasized by Management during a particularly impactful or long outage around 2+ years ago).
I haven’t studied enough business theory to know if there is a “positive hockey-stick” effect that will come into play shortly; i.e., while SmartThings has been working on improvement for years, it sure seems very, very slow to manifest. I guess there’s no reason that there won’t be a sudden upward bend in the reliability curve (due to a sudden successful critical-mass implementation of the factors you list).
Oh … I guess that’s also known as a “learning curve”. ![]()
At a time like this, it was interesting when I saw this thread (from March 27, 2015) yesterday while looking for something else…
To be clear, I’ve sort of already drank the kool-aid (especially since I can’t afford to change platforms any time in the near future), and I’m a really big fan. The problem now is that I’m starting to be a bigger fan of this Community forum and some of the SmartApp coders around here than the platform itself, or its makers (I know that sounded stupid and cheesy, but…).
I mean, I know the platform makes what these folks do possible in the first place, but seriously, where would we be without things like Rule Machine (@bravenel / @Mike_Maxwell), SmartTiles (@625alex / @tgauchat), and a bunch of other SmartApps and Device Handlers (most of which is on-hold while we wait for things to get cleared up).
I’m fine with waiting (betting a company as big and successful as Samsung will do what’s needed, and get this taken care of; especially since satisfied customers are their greatest assets, and they’re not a non-profit). I assume many aren’t.
My only question is this… How will we know?
How will we know when we’ve gotten to the point where the system is stable? I’m not just talking about the current system instability problems. I’m wondering about way down the road when the system has matured to the point that things like this can’t happen. How far down the road is that, and/or how will we know when we’re there? Will it simply be a matter of track record over time?
I appreciate the explanation from the engineer. However, it was very vague and most importantly did not answer whether or not SHM is working again or will work again anytime soon. Something happened at a specific point in time which cause all of us to have major issues of the same kind, you cannot tell me that is just a coincidence or a result of multiple factors. That simply does not make sense.
I feel like you are talking about it as if everything is better now, but if SHM still isn’t working I don’t understand how anything has changed at all.
It took too long, but the updates on the issues are appreciated. There’s no such thing as over communication for things like this.
It’s probably a good thing we can’t call and talk to someone live. Me, my neighbors and my gf have been aggravated these past few days. Thanks for something but it’s very vague.
Can I borrow this one?
The update is appreciated, while not extremely specific, I do not believe that it really is our place to have all of those specifics.
Either way, thank you for at least acknowledging this community. It really is a big step forward to a small group of users that probably push this system harder than ask if the rest combined.
I have seen a noticeable improvement this evening in my home system. I was excited that the system went into evening mode, 20 minutes before sunset local. Which is what it was supposed to do.
I’m not going to critique you’re work ethics or your practices. I do not know enough inside information on the company to even begin to try that.
What I will say is this, please, please absolutely learn from the mistakes that lead to this. Yes, there were mistakes, there there’s always is, as I do not believe in accidents out things that are unexpected. Especially in a highly monitored system.
Also, please continue, and please increase the participation of company employees into this forum. There is a lot to learn here. We may not be employees, but a group always had much more knowledge than a bunch of individuals.
Keep up the hard work, it will be noticed.
Thanks for joining the conversation, Benjamin; and for the highly detailed post. It is … nice to meet you.
In all honesty, though, I can’t say that is it particularly confidence inspiring. ![]()
Lots of platforms are “constantly growing and evolving”. I guess I can faintly remember the days that eBay had horrendous outages … and GMail, and probably a few other giants.
But that is the biggest concern of all. You see, the current rate of SmartThings growth is trivial compared to your goals; a couple hundred thousand hubs currently growing at, what, a few thousand a month; compared to a couple million hubs (growing exponentially) as Samsung UHD TVs come online and as marketing efforts suddenly reach critical mass and cause another “hockey-stick” of growth?
Your post did mention proactive risk management initiatives, and I appreciate your inclusion of them. I hope you and your team – and the entire company – follows through. Thanks!
Not true. I now use Rule Machine for almost everything and also have the Routines built in the special section for Routines. I can even manually run the routine that was built in Rule Machine by pushing them as well ![]()
I set everything up in Rule Machine first, then went back in and created the Routines in the special section, but didn’t selected anything to run when setting them back up. Went back to Rule Machine and created a new rule that says when a Routine button is pushed to run the Rule of what that Routine is supposed to do. Sounds confusing, but it works. I only set them up just in case a routine via Rule Machine didn’t run when it should have.
So is SHM working again or not? It was an extremely long-winded update that basically said nothing. And they keep saying they’re going to make improvements to make it better but how can you improve something when you admittedly don’t know what’s wrong with it?
I think most would agree with me when I say “please stop working on new features until the existing ones work consistently!”
That aside, SmartThings really needs to work on communication. If I weren’t in IT myself, I’m certain I’d never have found this great community. And without it, I would have never learned of status.smartthings.com, which frankly, I only learned of a couple of weeks ago (I’ve been using SmartThings since last August)!
And without knowing about that web site, what do you think I’d be doing as a SmartThings user right now? Yep, you got it–mashing every button in the app repeatedly and wondering why in the hell nothing is working (and likely making your problem worse as everyone is now increasing load on the system trying to figure things out). We get emails about hub upgrades, why not outages, especially major ones?
Good luck fixing things. I hope this is the beginning of the end of all the stability issues.
Please look into the aeon labs siren as an alert in SHM security. SHM security works again after removing the siren. Addng the siren back causes the arming/disarming to lock up.
Yes it was. This was all the rage (vaporware) when v2 was announced. Myself, and probably many many others preordered v2 in search of local execution and were promptly given the shaft when v2 was finally released.
Agreed. Please fix the Aeon Labs siren in SHM. Thanks.
Speaking of vaporware, remember the V1 to V2 migration tool they promised?
Also remember the things page? They took that away from us and told us what a wonderful thing they were gong to be replacing it with. Anyone seen that shiny new things page yet?
Or fix anything on that siren. I’ve never been able to change the volume or tone on the thing since I bought it. It just throws an error that I need to fill in all the fields, even though I clearly have. Support wasn’t any help and months later, the problem still exists. Turns out that if you change both values at once (even if you only want to change one) it seems to let you save your change then. Fun.
That’s leaving out all the times when the siren just goes off without any intrusion or when there is an intrusion and the siren doesn’t mute when you tap “mute alarms”. I yanked both SHM and the siren out last night and don’t intend on adding them again for a few days…