Azure Application Logs
If you host TomoriBot on an Azure VM, you can ship the bot’s structured error logs into an Azure Log Analytics workspace and build Grafana error panels on top of them. This page covers the whole pipeline: the bot-side JSONL file output, the Azure Monitor Agent ingestion, and the operational rules that keep ingestion reliable and cheap.
For monitoring a local instance with Grafana instead, see Local Grafana Monitoring.
Pipeline overview
Section titled “Pipeline overview”Pino structured logs (level >= 50 in production) -> /app/logs/tomoribot.jsonl in the container (TOMORI_LOG_FILE) -> /var/log/tomoribot/tomoribot.jsonl on the VM (bind mount) -> Azure Monitor Agent "Custom Text Logs" data source -> Data collection rule (DCR) parses each RawData line with parse_json() -> Log Analytics custom table (e.g. TomoriBotLogs_CL) -> Grafana error panelsThe DCR uses Custom Text Logs, not Custom JSON Logs. The file is UTF-8 JSONL, but
production error records contain nested err and context objects; collecting each line as a
single RawData string and parsing it inside the DCR transform is the reliable
Azure-supported route for nested JSON.
Step 1: Enable the JSONL file output
Section titled “Step 1: Enable the JSONL file output”Set TOMORI_LOG_FILE to a writable path inside the container:
TOMORI_LOG_FILE=/app/logs/tomoribot.jsonlBehavior:
- Only active in JSON output mode (production, or any run without the
pino-prettydev transport). Development pretty-printing is unchanged and ignores the variable. - Every emitted record is written to both stdout (so
docker logsdiagnostics keep working) and the file, as identical newline-delimited JSON, via an append-only synchronous stream. - Production logs at level
error(50) and above, which includes the custom levelsmetric(52) andrateLimit(55). An empty file on a healthy instance is normal.
Container wiring
Section titled “Container wiring”deploy/azure/docker-compose.yml sets the variable and bind-mounts the host directory over
the image’s /app/logs:
environment: TOMORI_LOG_FILE: /app/logs/tomoribot.jsonlvolumes: - /var/log/tomoribot:/app/logsThe container runs as the non-root user tomori (UID/GID 1001), so the host directory must
be owned by that ID. The Azure deploy workflow creates it before Compose starts:
sudo install -d -o 1001 -g 1001 -m 0750 /var/log/tomoribotVerifying the host file
Section titled “Verifying the host file”After a deploy, on the VM:
sudo ls -la /var/log/tomoribotsudo tail -n 3 /var/log/tomoribot/tomoribot.jsonlEvery nonempty line must be one valid JSON object. The file may legitimately be empty until the first production error occurs.
Step 2: Create the custom table
Section titled “Step 2: Create the custom table”Create a custom table (this guide uses TomoriBotLogs_CL; the _CL suffix is required) in
your Log Analytics workspace with the Log Analytics Tables REST API
(2021-12-01-preview), using this schema:
| Column | Type | Purpose |
|---|---|---|
TimeGenerated |
DateTime | Original Pino event time, with ingestion time as fallback |
Computer |
String | Populated by Azure Monitor Agent |
FilePath |
String | Populated by Azure Monitor Agent |
level |
Int | Pino numeric level |
msg |
String | Pino log message |
code |
String | Top-level or nested error code |
errorType |
String | Error type/name (type is a reserved column name and is rejected) |
commandName |
String | Command context when available |
message |
String | Error message with msg fallback |
RawData |
String | Original JSON line for access-controlled troubleshooting |
Wait until the table appears in the workspace before creating the DCR.
Step 3: Create and associate the DCR
Section titled “Step 3: Create and associate the DCR”Custom text log collection requires a data collection endpoint (DCE) in the same region; create one first and reference it from the DCR. Then create the DCR (platform Linux), associate it with the VM (installing Azure Monitor Agent if needed), and add a Custom Text Logs data source:
- File pattern:
/var/log/tomoribot/tomoribot.jsonl - Table:
TomoriBotLogs_CL - Record delimiter: end-of-line (correct for JSONL)
- Destination: Azure Monitor Logs -> your workspace
Apply this ingestion-time transformation. Its projected columns must exactly match the table
schema. The trailing where clauses drop known-noisy records (periodic cache metrics and
provider availability retries) to control ingestion cost:
source| extend p = parse_json(RawData)| extend EventTime = datetime(1970-01-01) + tolong(p["time"]) * 1ms| extend TimeGenerated = iff(isnull(EventTime), TimeGenerated, EventTime)| extend level = toint(p.level)| extend msg = tostring(p.msg)| extend code = iff(isempty(tostring(p.code)), tostring(p.err.code), tostring(p.code))| extend errorType = iff(isempty(tostring(p.type)), tostring(p.err.name), tostring(p.type))| extend message = iff(isempty(tostring(p.message)), iff(isempty(tostring(p.err.message)), tostring(p.msg), tostring(p.err.message)), tostring(p.message))| extend commandName = iff(isempty(tostring(p.commandName)), tostring(p.context.commandName), tostring(p.commandName))| where level >= 50| where msg != "metric:cache_sizes"| where RawData !contains "UNAVAILABLE"| where RawData !contains "RESOURCE_EXHAUSTED"| project TimeGenerated, Computer, FilePath, level, msg, code, errorType, message, commandName, RawDataDCR transformations compile against a restricted KQL subset, which shapes the query above:
coalesce() and unixtime_milliseconds_todatetime() are not available (hence the nested
iff(isempty(...)) fallbacks and the epoch arithmetic), and time is a reserved keyword in
the transform parser, so the Pino timestamp field must be accessed as p["time"].
Terraform ownership of VM attachments
Section titled “Terraform ownership of VM attachments”The existing workspace, DCE, DCR definitions, and custom tables remain external to Terraform, but
the production VM’s AzureMonitorLinuxAgent extension and its VM Insights/application-log DCR
associations are Terraform-managed. Their complete, non-sensitive existing resource IDs are set in
terraform/azure/terraform.ci.tfvars and must be updated if either DCR is deliberately replaced.
This boundary is intentional: replacing the VM deletes its extensions and associations, while the workspace and DCRs survive. The next Terraform apply therefore restores collection without rewriting the existing streams, table schemas, or ingestion transforms. Do not re-create the workspace or DCR definitions merely because a replacement VM has no incoming data.
Step 4: Test ingestion end to end
Section titled “Step 4: Test ingestion end to end”Append a harmless synthetic level-50 record to the host file (do not manufacture a real bot failure), wait roughly 5-10 minutes for initial ingestion, then query the workspace:
TomoriBotLogs_CL| where TimeGenerated > ago(30m)| order by TimeGenerated descIf nothing arrives, check: DCR-to-VM association, Azure Monitor Agent status on the VM, exact file path and directory permissions, that each line is valid one-line UTF-8 JSON, DCR error metrics, and that the transform’s projected columns match the table schema exactly.
Step 5: Cache-size metrics (second dataflow)
Section titled “Step 5: Cache-size metrics (second dataflow)”The logger emits a periodic cache_sizes sample (custom level 52, one line every
CACHE_METRICS_INTERVAL_MS, default 5 min) carrying every in-memory cache count plus process
RSS. The error dataflow above deliberately drops these (msg != "metric:cache_sizes") so the
error table stays clean. To chart them, the same DCR carries a second dataflow that reads
the same file stream but keeps only cache samples and projects their numeric fields into a
dedicated table, TomoriBotCacheMetrics_CL (typed int/real columns for each cache plus
rss_mb/rss_pct). Two dataflows over one input stream is the supported pattern; their
where filters are disjoint, so no record is double-counted.
Second dataflow transform (abridged — one toint(p.<field>) per cache column):
source| where RawData contains "metric:cache_sizes"| extend p = parse_json(RawData)| extend TimeGenerated = datetime(1970-01-01) + tolong(p["time"]) * 1ms| extend rss_mb = toreal(p.rss_mb), rss_pct = toreal(p.rss_pct)| extend shortTermMemory = toint(p.shortTermMemory), tomoriState = toint(p.tomoriState) // ...| extend discord_users = toint(p.discord_users), discord_members = toint(p.discord_members) // ...| project TimeGenerated, Computer, rss_mb, rss_pct, /* app caches */, /* discord.js caches */Cost note: this reverses the original “drop cache_sizes to save ingestion” decision, but the
volume is tiny — ~288 samples/day at ~900 bytes (~260 KB/day), well within the free tier. The
discord.js caches (discord_users, discord_members, discord_channels, discord_presences)
are the real memory story; the app caches are usually small.
Grafana
Section titled “Grafana”Point panels at an Azure Monitor data source whose identity has Reader on the resource group
and Log Analytics Reader on the workspace. That identity needs read access only —
ingestion permissions belong to the DCR/agent path, never to Grafana. Current panels:
- Error count (
Time series) —TomoriBotLogs_CLcounted per minute.make-seriesreturns packed arrays, so append| mv-expand TimeGenerated to typeof(datetime), Errors to typeof(long)or the series will not render. - Error log (
Logs) —TomoriBotLogs_CLshowingTimeGenerated, a title fromcode/errorType+message,commandName, andRawData. - Cache sizes (
Time series, stacked) —TomoriBotCacheMetrics_CL, one series per cache column, for spotting growth/leaks. - Token consumption (
Time series, stacked bars) — Postgresstat_counters,SUM(count)oftokens_in+tokens_outgrouped bybucket(day) andmetric_key(model). Daily grain only; the table stores no finer bucket.
The token panel uses Grafana’s separate read-only PostgreSQL login over TLS verify-full; it must
not use the application runtime or administrator credential. Keep the Azure Monitor VM/error/cache
panels alongside the direct PostgreSQL statistics panels. The retired
Custom-TomoriBotTokenMetrics_CL DCR flow must not be recreated: PostgreSQL remains authoritative
for token, cost, and usage statistics.
Beware GCP-carryover panel state: legend-click “hide series” overrides and series-name color
rules reference the old query’s series names and silently misfire on Azure queries (e.g. a
byNames: ["tomoribot"] exclude override blanks an Azure series named value).
Operational rules
Section titled “Operational rules”-
RawDatais privileged troubleshooting data. Access is limited to identities with Log Analytics read access, currently the designated operator/Grafana identity. It can contain error context after structured redaction, so do not publish it in GitHub Actions output or widen workspace roles. Retain it while nested context is needed; reconsider removing it only after parsed columns have proven sufficient. -
Retention is 30 days. Keep the workspace at the production 30-day retention baseline unless an explicit incident-response or cost decision changes it.
-
Keep the file append-only. Do not add log rotation that renames files into a pattern the DCR’s file glob still matches — Azure Monitor Agent treats a renamed file as new content and duplicates ingestion. Design retention so rotated files fall outside the collected pattern, and only after the pipeline is verified.
-
Cost control lives in the DCR transform. Prefer adding
whereclauses there over widening what the bot writes; everything that passes the transform is billed ingestion. -
stdout is unaffected.
docker logskeeps working regardless ofTOMORI_LOG_FILE, so the file can be wiped or the variable removed without losing container diagnostics.