Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.
Use service, wait for 404 or data loss…
Submitted 1 day ago by
user224@lemmy.sdf.org@lemmy.sdf.org to selfhosted@lemmy.world
Running services, network usage, memory usage, bandwidth, disk I/O, successful logins, whether the thing is even alive, etc…
So far my only method has been “hope and pray”.
Use service, wait for 404 or data loss…
I notice something stopped working, or someone in the household notices.
Then I go fix it.
DeathByDenim@lemmy.world 11 hours ago And the best time to notice that the backups weren’t working is when you need access to the backup. 😅
Backups are pretty much the only thing I get notifications for. Otherwise, one app or another on my phone will complain if it loses connection to my server.
Backups? In this economy?
timochka@lemmy.zip 3 hours ago I use Signoz and OTel collectors to bring in everything from Kubernetes (and also SNMP-to-OTel for the network and NAS stats.)
That covers dashboards and basic alerting. Then on top of that I have an n8n workflow that runs every morning - it pulls all the main stats out of Signoz, as well as issues and their status from my infrastructure repo in Gitlab, and feeds it all into a local LLM (Qwen-3.8-Flash-Next at the moment), which emails me a report on overall cluster health. It also automatically opens a ticket in Gitlab for anything that might need fixing.
It works really nicely, gotta say.
owenfromcanada@lemmy.ca 1 day ago I have a robust monitoring system for my Jellyfin server, been running for the last few years. Checks in periodically, at least once per 24 hours and notifies me if it’s down. Doesn’t use any electricity, but does consume a good amount of Cheerios and mac & cheese.
Does this monitoring system also happen to be a dependent on your taxes?
I’ve got a few of those services running.
Lies! It uses electricity!
owenfromcanada@lemmy.ca 23 hours ago Indirectly. But as another commenter pointed out, it balances out with the tax rebates.
Gatus. Works and easy incremental set up.
kurcatovium@piefed.social 18 hours ago That’s the neat part, I don’t.
Lights are on, server is on.
I recently found out that this is not always true.
I never said anything about it responding.
Kernel panic?
My kernel is chill.
Beszel & Pangolin.
Tealk@rollenspiel.forum 6 hours ago Icinga and Phare Uptime
Whining from end users.
I pay a group of schoolchildren to refresh a group of browser tabs and yell if anything isn’t responding.
There are plenty of previous threads with recent answers if you search.
Ultime Kuma and Beszel
Well I used to use VGA for my monitors, now its mostly displayport.
Seriously though, prom+Loki+alloy+a few other things to Grafana, with alerting in grafana.
It’s running the private DNS for my phone. If it goes down, I realise quite quickly.
Uptime Kuma! For HDDs I have them report their “ping” as percentage full
Can you teach me this wizardry?
Sure! In Uptime Kuma, add a monitor with type Push. It gives you a URL with a unique token, and the monitor goes into “down” state if nothing calls that URL within the heartbeat interval. Then have a cron (or a systemd timer or whatever) run this on the machine with the drive every 5 or 10 minutes:
use=$(df --output=pcent /mnt/yourdrive | tail -1 | tr -dc ‘0-9’) curl -fsS -G “http://your-kuma:3001/api/push/TOKEN --data-urlencode “status=up” --data-urlencode “msg=${use}% used” --data-urlencode “ping=${use}”
Ping is of course meant to be response time, but Kuma will graph any number you toss in, so you get a “filled” chart for the drive. I also send status=down once usage passes a threshold. And if the box itself should fall over for some reason, the missing check-ins trigger the alert too.
(I typed this quick on my phone, so please double-check and probably don’t blindly copy-paste)
Monit, simple to setup, gets the job done, runs on my router too.
I have this on my NAS, works well
moldy_rice@piefed.keyboardvagabond.com 10 hours ago I go the easy route with munin and wazuh. It comes with a ton of stuff preconfigured
My primary monitoring system is Beszel, but since I just migrated from VMs to LXCs it can’t properly see the memory/CPU usage, so it just checks for disk usage and liveliness. Home assistant with Proxmox & Dockhand integrations gives me the rest of my monitoring needs.
I’ve been checking several other solutions before and Beszel just hits the spot.
If you run Proxmox and Proxmox Backup Server, I actually JUST found a better product for my needs, Pulse
My only problem with Beszel is that it doesn’t know about LXCs within Proxmox, so I can’t see if there are CPU/memory issues happening within them. Pulse integrates directly and can see all of it!
If you don’t run LXCs within Proxmox, I’d say stick with Beszel.
I am also using Beszel after using checkmk. For me checkmk was way too complicated (but is also way more powerfull) to setup and maintain propperly. So I switcheed to beszel because it is much easier and good enough for my use case which is:
Besides that I don’t need anything. The only thing missing is SNMP support to check switches, routers and AP’s.
Kolanaki@pawb.social 10 hours ago I just look at the console.
I am using prometheus-nodeexporter which is scraped by Prometheus and then visualized in Grafana
Grafana dashboard
image You can find more details and the source code here: https://erasmus.works
user224@lemmy.sdf.org@lemmy.sdf.org 16 hours ago Ooooo, that looks nice.
I use naemon (nagios fork) and grafana Loki Prometheus influxdb
happy_wheels@lemmy.blahaj.zone 1 day ago Cockpit. Total gamechanger for me. I don’t have to ssh to my boxes individually to see what’s going on. Instead it’s a web UI AND you can connect basically every server you want to one instance.
github.com/cockpit-project/cockpit
That said. For individual servers?
htop/nmon (nmon is FANTASTIC)
ps -ef | grep <whatever you’re looking for>
dmseg
tail -f /var/log (or whatever log file you care about)
netstat
those are all that come to mind right away.
It’s amazing the risk profile of that.
How do you configure adding multiple servers to the same instance?
happy_wheels@lemmy.blahaj.zone 12 hours ago So I actually just went to make a guide on this, demonstrating on my ubuntu 26.04 lab vps, and it appears the connect option is not there.
However on my 24.04 lab vps, it’s there. Same with my local 22.04 systems (yes I know I have to upgrade lmao)
I went to cockpit docks and sure enough, the feature is deprecated :( Dang.
docs.cockpit-project.org/…/feature-machines.html
It was really cool though while it worked.
Lol I go to all of my server’s specific cockpit addresses, sounds like I’m doing it wrong, I’ll need to look at connecting them
happy_wheels@lemmy.blahaj.zone 12 hours ago So it turns out… you can’t connect them anymore… The feature is deprecated
docs.cockpit-project.org/…/feature-machines.html
(also, I wanted to just link my response from another user who had asimilar question but for whateVER reasyon, Alx won’t let me??)
popekingjoe@lemmy.world 20 hours ago I usually just walk down the hall and move the mouse. I don’t really need remote monitoring. 😅
Prometheus + grafana. It’s overkill tbh, most of the services restart automatically and I mostly ignore it. (it’s an artifact from life where I cared about it)
Exactly this scenario at my place as well. 😄
I use Alloy to collect Metrics of the host (Disk usage, CPU, Ram, etc.) and different logs. With Grafana the data is then displayed as a daschboard. If everything goes south, Alertmanager sends Mails to me
Acronyms, initialisms, abbreviations, contractions, and other phrases which expand to something larger, that I’ve seen in this thread:
| Fewer Letters | More Letters |
|---|---|
| AP | WiFi Access Point |
| DNS | Domain Name Service/System |
| LXC | Linux Containers |
[Thread #122 for this comm, first seen 9th Oct 2026, 13:00] [FAQ] [Full list] [Contact] [Source code]
I get a text pretty quickly.
Uptime Kuma for outages, Prometheus for metrics, rsyslog to Loki backend for logs. Grafana for ingesting Prometheus and Loki data.
bertlewirth@lemmy.world 12 hours ago
You guys are monitoring your servers?
InFerNo@lemmy.ml 12 hours ago
“can’t access this thing, server must be down”
TrollTrollrolllol@lemmy.world 6 hours ago
I’ll fix it when I get home
I plan to perhaps, just maybe, I am not sure, to try self-hosting an email server. But I’ll need a good uptime.
Although I am really scared about security.
I self-host my own email, have done so for years. I’ve not gotten pwned. Use an SSH key and a firewall, only expose the ports that should be public, have reasonable access controls, etc, just practise general cybersecurity—it applies to servers too.
I know everyone complains about hosting email. But I’ve not run into major problems. My emails get through spam filters and I have good uptime. It’s also a fun learning experience.
bertlewirth@lemmy.world 12 hours ago
Well, it’s a good learning experience. But…… don’t. Good luck getting around the major email providers’ smart screen shit. It’s probably not gonna happen. I tried a few years ago, and even microslop couldn’t figure out why their shit screen crap wouldn’t allow my emails. But it’s worth doing it in the short term to learn how to do it.
InFerNo@lemmy.ml 11 hours ago
Been running one for many years, mostly issue-free. Had to request whitelists way in the beginning, but things turned out ok.
Check out this super great guide. workaround.org/ispmail-trixie