System/Network/Application Monitoring for a full, high-performance stack -- What do you use?

Events happening in the community are now at Drupal community events on www.drupal.org.
_-.'s picture

I've installed Zabbix 1.8.3 monitoring on my Drupal (actually, Pressflow) box.

It works great out-of-the-box for non-Drupal-specific server-level monitoring, especially across multiple/distributed servers.

Now, with Drupal/Zabbix module (http://drupal.org/project/zabbix), it works nicely for monitoring Drupal Core events. At least when 'just' Drupal listens directly at :80.

But, with a more 'convoluted', high perf stack (e.g., multi-site Pressflow/Apache, PHP+APC, Varnish cache/proxy, Memcached, Nginx as edge server, etc etc), representative Zabbix monitoring seems to require a bit more 'art'.

I suppose I could point probes/monitors @ each of the front-end server(nginx), cache (varnish), & backend (apache) HTTP listener ports. I think that'd be sufficient to get me at the 'up/down' state info. All the 'additional details' -- cache utilization, hit rate, errors, etc -- would probably need to have custom written probes.

Zabbix certainly has the capability to extend/customize probes for just about anything ... Overall, it's a highly recommended package, and can be commercially-supported if/as req'd. But, afaict so far, none/few of the probes required for the "Pressflow stack" are already available to plug in.

Some I've chatted with use collections of different utilities -- e.g., nagios+munin/cacti+smokeping/pingdom etc -- for a "full-stack" monitoring solution.

@ Munin community I found some varnish plugins available, which are also importable into Zabbix (http://www.zabbix.com/wiki/?do=search&id=munin), but they seem to currently have a few issues ...

What do you use for a 'full stack' solution?

Given the increasing #s of people running these 'convoluted' high-performance stacks, perhaps there'd be interest in cobbling together a portfolio of relevant probes & info for consolidated monitoring.

"Here" seems as good a place as any ...

Comments

interesting subj, subscribing

wik's picture

interesting subj, subscribing :)

fyi, its seems that one of

_-.'s picture

fyi, its seems that one of the providers of 'the complex stack', Pantheon, is at least interested in similar issues,

Short Term Pantheon Roadmap
http://groups.drupal.org/node/58853
"Monitoring: We're evaluating a lot of possible answers here, ..."

I've not yet found further info on what they're doing (if you have, a link?), or thought much yet about if there's a one-size-fits-all, best-practices solution that'll come out of it.

My 1st inclination is to suggest that a general set of approaches/tools/plugins, that can be then focussed on specific stacks (like Pantheon/Mercury) is an approach with broader appeal to the wider community here.

Write custom probes for Zabbix

agileware's picture

I suppose I could point probes/monitors @ each of the front-end server(nginx), cache (varnish), & backend (apache) HTTP listener ports. I think that'd be sufficient to get me at the 'up/down' state info. All the 'additional details' -- cache utilization, hit rate, errors, etc -- would probably need to have custom written probes.

I think you've hit the nail on the head. If you want more detailed statistics etc. from other systems comprising the 'stack' then just write custom probes that report back to Zabbix using the local Zabbix agent. That's how I would do it.

Agileware are an Australian Web Team specialising in Drupal, WordPress and CiviCRM.
Agileware provide development, support and hosting services.
https://agileware.com.au

A hosted 'library', somewhere

_-.'s picture

A hosted 'library', somewhere in Drupal-land, with documentation, of such custom-written probes & plugins for various components common to the stacks typically used 'here', with implementations for the various 'big' players, would be a nice-to-have.

I'm guessing that numerous folks have implemented bits -n- pieces already. Sharing them centrally would help prevent reinventing the wheel the n-th time, and would certainly save on time for endless sleuthing to find all the pieces ...

We like ScoutApp and Droptor

jemond's picture

We like ScoutApp, which is a hosted server level monitoring solution. It doesn't give you quite the same level of network level monitoring that a Cacti/Munin might, but it gives us the data we are most interested in (load, memory usage, peaks, etc) and is dead easy to install on a Linux box.

For monitoring at the Drupal-level in the stack I want to mention our tool called Droptor:
http://drupal.org/project/droptor
http://www.droptor.com/tour

We built it to specifically provide monitoring at the Drupal layer of a stack. Memory profiling was added to Droptor 2.0, which launches tomorrow:
http://www.droptor.com/v2

subscribe.. interesting

J0keR's picture

subscribe.. interesting

Cacti

rjbrown99's picture

I've been using Cacti as my primary trending tool, which has great support for monitoring the system, IO, daemons, etc. What it was lacking was the ability to directly monitor Drupal. To that end, I started http://drupal.org/project/cacti. Right now it's more of a framework since it really only has one graph of authenticated vs anonymous users. The idea is to add a hookable framework for any module to expose data in a Cacti-supported format. It can then be ingested into Cacti with the right graph/data template and plotted on the same timeline as the actual system performance data. I'd love to see some additional users, issues, and feature requests, patches, etc.

On another note, I send all Drupal logs to syslog and use Splunk - http://www.splunk.com - as the main interface. It's STUPIDLY easy to set up, on an RHEL-based system it was under 5 minutes. The interface is beautiful and it allows you with pretty much zero training to start doing intelligent queries on all of your log files. I can't say enough good things about it.

High performance

Group notifications

This group offers an RSS feed. Or subscribe to these personalized, sitewide feeds: