NRAMP Docs
← Sign in
Documentation

NRAMP Docs

What each module answers, how to read it, and what it costs to run. This is the same help the Help button shows inside NRAMP, for every screen in one place.

Getting started

Overview

NRAMP runs as a service on a private server of ours, reached in a browser over HTTPS. It is not public: there is no sign-up, accounts are created by an admin here, and every page needs a session. It talks to New Relic with the API key an admin sets in Admin, and keeps everything it collects in its own database on that server. Nothing goes anywhere else, and the New Relic key itself never reaches your browser.

Each card opens one task.

Every module is listed whether or not your account has it. A module you have not been granted is greyed, its badge reads No access instead of Available, and clicking it says so rather than opening anything. That is deliberate: a hidden card makes the launcher a different shape for everybody, and the answer to "can I get at X" becomes that X does not appear to exist. Seeing it lets you ask an admin for it by name, and it grants nothing, since the server refuses the paths behind it either way.

The modules

  • Ingest Analysis. Track every account's ingest against its own history, spot trends, identify contributors to trends, and recommend remediation.
  • CCU Analysis. Snapshot of heaviest CCU usage: people, alert conditions and dashboards consuming the most CCU across every account, and their trends.
  • Alert Analysis. Audit any account's alert policies and conditions in full, what can fire, what is silenced, what is broken, what is not needed.
  • Incident Heatmap. 24 hour to 7 day incident heatmap across the estate, the individual account or condition.
  • Dashboard / Query Profiler. Browse an account's dashboards, identify which queries cost the most compute.
  • Log Partition Analysis. Analyze log partitions and identify recommended partitions. Review the impact of implementing them.
  • Log Validator. Review an account's log data and identify logs ingested but not used.
  • Metrics Validator. Review an account's metric data and identify metrics ingested but not used.
  • Ingest Budgets. Build budget plans and track account utilization against monthly budgets.
  • Account Assessment. Create an account summary of all recommendations and findings from NRAMP.
  • Migration Tools. Migrate New Relic entities (dashboards, alerts, synthetics and workloads) between accounts, alone or with others.
  • Chargeback. Attribute ingest cost back to the teams and business units that generate it, so the bill lands with the people who can act on it.
  • Workstreams. Track the NRAMP recommendations tasked for implementation in workstreams and the related tasks.
  • User Audit. Identify a user's access, and recent activities within New Relic.
  • Environment As-built. Identify what is deployed and configured across all accounts.

The NRAMP mark in the top left returns here from anywhere, on every screen, including the two that open in a tab of their own. Back, beside it, returns to the page you came from rather than here: Ingest Analysis to Admin and back lands on Ingest Analysis, and a workstream returns to its list. My Links, at the bottom, holds shortcuts you keep for yourself.

Your account

Your name in the top right opens your account. Change password asks for your current password and the new one twice; saving it signs you out of every other browser and keeps you signed in on this one. After five wrong current passwords it waits before letting you try again. Theme, beneath it, switches between light and dark for everyone, whatever their role. The choice is remembered in the browser you are using, so it follows the machine and not your account, and picking the theme already in use hands it back to the operating system's own setting. The same window holds your API keys, which call NRAMP's own API with your access and no more. A session lasts twelve hours, so expect to sign in about once a day.

Ingest and cost

Ingest Analysis

One row per account, showing ingest for the current partial month plus the five complete months behind it, as monthly totals and as daily averages. The filter box above the table matches an account's name or id, and several terms separated by commas show accounts matching any of them, so "GTIO, BRE" lists both.

Account or budget

The toggle beside the filter decides what the table lists against. Account is the default and the table as it has always been. Budget adds two columns after the name, the Owner and the GBL from the Chargeback table, and opens sorted by GBL so the estate reads as the budgets that pay for it rather than as a list of accounts. Accounts not in the Chargeback table say so in the Owner column and sink to the bottom of that sort.

Clicking an Owner or a GBL that has a value opens the Chargeback screen with an invoice already filled in: every account sharing that owner or that GBL, this month, titled after the account clicked from. The invoice is generated but not saved, so it can be read, adjusted and printed, and saving it stays a deliberate press. Accounts that reported no ingest this month cannot be billed and are not in the picker at all, so the line under the button says how many were left out.

Both columns sort like any other, the filter and Alerts only still apply, and switching back to Account removes them and returns the table to its usual order. The toggle appears only for people granted Chargeback, because who pays for an account is that module's data; your choice is remembered.

Ingest against budget

Beside the title: an arrow, the month so far, where it is projected to land, and the budget set in Admin. Projected is the month so far plus a daily rate for the days still to come. From a week into the month that rate is this month's own daily average, which makes the projection that average held all month. In the first week the rate leans on last month's average and moves to this month's as the days accumulate, because on the first of a month the "daily average" is a few hours scaled up to a day, and thirty days of an evening's traffic is not a projection. It and the month to date both turn red when that projection lands above budget and green when it lands below.

It is colored on the projection rather than on the figure so far, because a month is always under budget on its first morning and saying so would tell you nothing. Hover it for the daily rate and where the month is heading. The arrow is this month's daily average against last month's, using the same colors and the same flat band as the trend columns in the table. For a month's first three days there is no arrow, here or in the daily averages column: a dot stands in its place and says how old the month is, since a comparison made from a few hours points wherever those hours happened to. With no budget set, the figure is shown without a verdict.

Why daily averages matter

The current month is partial. On the third of the month it holds three days of data, so comparing its total against a full month would make every account look like it collapsed. Dividing each month by its elapsed days makes any two months comparable on any date, and everything the app decides is based on that figure.

Reading the colors

  • Green. Below the trending line.
  • Yellow. Trending. At or past the trending fraction of the threshold.
  • Red. The variance threshold was met or exceeded.
  • Gray. No data reported, or no baseline to compare against.

Gray is deliberately not green. An account that stopped reporting has a problem, and showing it as healthy would be the more dangerous default.

The dot in the variance column

For the current month only, a variance below the trending line is calculated but not shown. Early in a month the daily average rests on a few hours of data and swings wildly, so a precise looking number there would be mostly noise. Hover the dot to see what it is actually running at. Complete months always show their variance.

Trend arrows

The two arrows answer different questions, and each sits beside the columns it is drawn from.

The monthly totals arrow compares the last complete month against the one before it. A partial month's total is smaller for no reason other than the date, so it is left out.

The daily averages arrow compares the month in progress against the last complete month, which is what the figure at the top of the screen does. A daily average is not distorted by a partial month, so waiting for the month to end would report last month's direction while the columns beside it already show this month's. Hover either arrow for the two months it used and how many days each covers.

Click any account to load its data sources below. Sizes are scaled to real units, so a small account reads in MB or KB rather than as 0.00. Hovering a figure gives the exact value in GB.

Printing an account

With an account selected, Account Detail offers two printouts. Print table prints the data sources table exactly as it is on screen: the same columns, the same sort order, and GB or dollars as chosen. Print report is a summary of the account: this month against last, the variance against its threshold, the month by month history, and every data source with its status and share of the month.

Saved as a PDF, either one is named NRAMP followed by the account name, for example NRAMP-Aurora Retail Prod GCP.

Ingest and cost

Variance analysis

This page explains why one data source grew. It compares a recent window against the same length of time 28 days earlier and attributes the change to individual contributors: metric names, log sources, applications, hosts.

Print page, in the header, puts the whole analysis on paper or into a PDF. The controls come off, the sheet is stamped with the date it was taken, and the per-contributor drop rules are opened so the printed copy carries each rule rather than the summary that hides them.

Why 28 days and not 30

New Relic keeps logs for 30 days. At exactly that edge the function used to size data stops returning figures while the event count still answers, so a 30 day baseline reads as zero bytes on every log source and makes everything look newly enabled. Sitting at 28 days keeps the comparison inside retention.

Three kinds of empty baseline

  • Measured. Normal. The percentage is meaningful.
  • Not measurable. Events exist in the baseline window but cannot be sized. The page says so and withholds the change, the percentage, and the NEW flags.
  • Genuinely absent. The source had not started reporting yet. The page names the earliest record it can find and explains that everything is new rather than grown.

Coverage

Contributors are split by whichever attribute best explains the volume, not merely the first one present. An attribute carried by a fraction of records would split neatly while leaving most of the bytes unaccounted for. The page reports what share the split actually covers.

Remediation impact

Actions are ranked by what each recovers, each with a risk level. The totals are a union rather than a sum, because actions overlap and adding them would claim more recoverable volume than the source ingests. Each action can become a task.

### Drop rules

Below the actions the page gathers every contributor those actions named and writes them out as pipeline cloud rules. One combined rule drops all of them at once, and One rule per contributor opens a separate rule for each, so you can apply them one at a time or hold back the ones you are unsure of. Each rule carries what it saves per day and per month, and a contributor that belongs to an action asking you to confirm the telemetry was meant to be on is marked verify first. Copy puts a rule on the clipboard. Drop rules are not retroactive and dropped data cannot be recovered, so run the equivalent SELECT first.

Results are cached for 24 hours. Past that, the stored figures are still shown straight away and a fresh scan runs behind the page, because scanning a very large source takes minutes; open the screen again shortly for the new ones. Re-run waits for a fresh scan instead.

The largest sources cannot be scanned whole inside the time New Relic allows for a query: Aurora Retail Prod logs about 20 TB a day, and a faceted scan of it runs right up against that limit. Where that happens the screen measures the most recent part of each window and scales it up, and says so.

Ingest and cost

Log record size

This page finds the log sources whose records are unusually large. It measures the average size of a record in one log partition, then compares every source against it and flags those more than the chosen percentage above.

What size means here

Sizes are what New Relic bills, which counts every attribute on a record and not just the message text. In practice the message is often a small part of a record. The Attributes per record column shows the rest: a source far above average with an ordinary message length is carrying extra attributes, and the fix is to drop attributes rather than shorten what it logs.

Sources and records are measured differently

The billed size can only be measured over groups of records, never one at a time. So sources are compared on their true average, but the largest individual records inside a source are found by message length, which is the only per-record size New Relic can filter on.

Printing

Print page puts the whole measurement on paper or into a PDF: the cards, the finding, and every source with its columns. The controls come off, the sheet is stamped with the date it was taken along with the partition and window it covers, and the table is set smaller so the last columns land on the page rather than falling off the right of it.

Bytes or dollars

The toggle beside Analyze shows the figures at the ingest price set in Admin, and remembers your choice along with the Accounts page. Daily volumes become dollars per day. A single record costs a tiny fraction of a cent, so per-record sizes are shown as the cost of a million such records instead, which keeps sources comparable. Message length stays in characters. Without a price in Admin, dollars are unavailable.

Excess GB per day

How much ingest would disappear if a source's records were the partition's average size, scaled to a day. Sources are listed in that order, so the first rows are where trimming pays most, not merely the largest records.

Message patterns

Opening a source groups its messages by pattern: the leading timestamp, level and thread are skipped, and the text stops at the first digit, so lines that differ only in ids and numbers fall together.

Tokens in logs

Records that look like they contain access tokens (a JWT, or a bearer header) are counted and marked. Anyone with access to those logs can read them. Text shown on this page has tokens, credentials and email addresses removed before it leaves the server; the logs in New Relic still contain them, and nothing here is stored.

Tasks

Every finding can become a task: a source flagged as large, a source whose records carry tokens, and, inside a source, a large or token-bearing message pattern. Task buttons appear only for people whose account may reach Workstreams; for anyone else the finding is simply shown. The task opens with the measurements, the fix most likely to help and a query that reproduces the numbers in New Relic. Log text in a task has had its tokens removed, like everything else on this page. A finding can live in one workstream only, and once tracked it says so here, whatever threshold the page is run at. Creating tasks needs Workstreams access.

Group sources by

Which attribute identifies a log source varies by account. Only attributes with more than one value in the last hour are offered, since a single value has nothing to compare against.

Ingest and cost

CCU Analysis

Where compute is going. Three rankings across every account you can see: the dashboards, the people, and the alert conditions burning the most CCU.

Every row leads to the screen that explains it. A dashboard opens in the Dashboard Profiler, on that account, with its widget queries and what each costs. A condition opens the Incident Heatmap for its account, filtered to that condition, hour by hour. A person opens User Audit for their address.

A condition's CCU is what evaluating it costs, which it pays every cycle whether or not it ever breaches, so the most expensive condition in an account can be one that has never fired. The Fired column counts the incidents each condition raised in the last 30 days, and marks the ones that raised none, which is where that money is most clearly wasted: seventeen of Aurora Retail Prod's fifty dearest conditions have never fired, the first of them costing over nine thousand CCU a day. It is read when the screen opens rather than collected with the snapshot, so a snapshot taken before the column existed still fills it in, and it is kept for half an hour because a thirty-day count barely moves.

A condition sent from here that is not among the fifty gets a row of its own on the map, named and linked in the first column like any other, carrying the hours it actually has. One that never fired has a row of empty cells and is marked as such, which is the honest shape of that answer; the note above the table says why a condition can cost a great deal and fire nothing. The name and its link come from the entity catalog rather than from the incidents, since a condition that never fired appears in no incident at all, and where an account runs two conditions under one name the policy the row came from picks the right one.

What is counted

CCU is New Relic's compute unit, separate from ingest. The figure is the billable CCU metric only. New Relic also reports AdvancedCCU and CoreCCU, which are the two halves of that same total, so adding all three would count everything twice.

Why collection takes a few minutes

Compute does not roll up the way ingest does. The rollup account reports only its own compute, so this queries every account in turn. A snapshot is kept for 24 hours, since compute is a daily figure and a day-old view is already stale. Re-collect ignores that.

Trend

Each row is matched to the same dashboard, person or condition in the previous snapshot, by name rather than by rank, since rank moves on its own. Up is shown in red because more compute is more cost. A row with no match is marked new rather than up from zero, which would read as a spike that never happened. With only one snapshot stored there is nothing to compare and the column reads n/a.

The snapshot date

It sits in each panel header rather than in every row, because it is a property of the snapshot and would otherwise repeat identically down fifty rows of three tables.

Coverage

Only compute that names its consumer can be ranked. Alert evaluation names its condition and policy, and dashboards and interactive queries name theirs. Background processing such as pipeline rules is real compute and is counted in the totals at the top, but it has no dashboard or person to attribute it to, so it appears there and not in the three tables.

Ingest and cost

Ingest Budgets

What an account is expected to use in a month, and whether this month's rate will take it there.

A plan is a tier, not a pot

A budget plan says what one account in it should use in a month, so "Large market, 100 TB" is a size several accounts can each be held to. The budget is not divided between them and their usage is not added up against it.

That is what makes the two findings the screen exists for possible. An account heading over its tier has ingest worth remediating. An account heading well under it is usually in the wrong tier, and moving it down is the cheaper fix.

The four thresholds are percentages of the budget, above and below, because one plan has to fit whatever account is put in it. A critical line has to sit outside its warning line, or the warning could never fire first and one of the two means nothing; the screen refuses a plan written that way.

What the color means

Health is judged on where the month will land: what the account has ingested so far, plus a daily rate for the days still to come. From a week in, that rate is this month's own daily average; in the first week it leans on last month's and moves across as the month fills in, and the line under the heading says when that is happening. The events sent to New Relic carry the same projection, so the first of a month does not raise breaches from a few hours of traffic.

A month-to-date figure would read green on the 2nd and red on the 30th whatever the account was doing, which is no use for deciding anything. A daily average does not care what day it is, which is the same reason the accounts screen leads with one.

Green is inside both warning lines. Yellow is past a warning line, over or under. Red is past a critical line. An account the latest snapshot does not carry is grey rather than green: not measured is not the same as on budget.

The two views

Plans is the default: one card per plan with its name, its per-account budget, how many accounts are in it, and a bar showing how those accounts divide between the states. A plan's color is its worst account, because a tier holding one account 90% over is not a healthy tier however many others are fine.

Click a plan for the accounts inside it, with month to date, daily rate, projection and how far each sits from the budget. Clicking an account card opens the same list with that account marked, so it can be read against the others in its tier.

Each row carries an Ingest button to that account on the Accounts screen, with its data sources and what each of them is doing. The budget says an account is over; the next question is always which source, and that is the screen that answers it. The button appears for readers who have that screen.

Accounts switches the cards to one per account, each showing its plan, its budget and its own indicator.

The second toggle beside it sorts by name or by size, and its labels follow what is on screen: plan name and plan size over the plans, account name and account size over the accounts. Size is the default, because the question that brings somebody here is usually which of these is big; name is how you find one you already have in mind. The bar's middle is the budget, not empty: being far under is a finding too, and a bar that only grew to the right could not show it.

Who can change what

Anyone granted the module can read both views and open a plan. Creating, editing and deleting plans, and moving accounts between them, is an admin's, because a budget is what everybody else is measured against. The server enforces it as well as the screen.

Assigning shows every account as a ticked list rather than a dropdown, with a filter and a box for narrowing to what you have ticked, so several accounts can go into a plan in one go.

An account can only ever be in one plan. Ticking accounts that are already in one is allowed and moves them, and the dialog says how many of your ticks that applies to before you assign.

A plan can only be deleted once it is empty. One that still holds accounts is refused, with the count and what to do about it, and the accounts open so they can be moved: deleting it would quietly un-budget every account in it, and emptying it first is the same work done where it can be seen.

The account list when assigning is alphabetical, because it is read by somebody looking for a name they already have in mind. The accounts screen orders by ingest, which is the right order there and no help at all for finding one account among two hundred.

Telling New Relic about a breach

A plan can send its warnings and criticals back into New Relic as custom events, which is how a budget becomes something that pages somebody rather than a color on a card you have to remember to look at. Tick Send a custom event to New Relic for warnings and criticals on the plan, and a bell appears on its card.

The event type is NRAMP_budgets. It is written with an underscore because New Relic accepts only letters, digits, underscores and colons in an event type and rejects the whole batch over a hyphen.

One event goes out per breaching account, carrying its status (Warning or Critical), which side it breached on (High or Low), the account name and id, the plan and its budget, the threshold percentage that was actually crossed, the budget owner, GBL and cost center from Chargeback, and the month to date, daily rate and projection in GB. An account inside both warning lines sends nothing: on budget is not news.

Alert on them in the events account with a condition over SELECT count(*) FROM NRAMP_budgets, faceted by accountName, filtered to the status you care about. The threshold is on the event, so a condition can say "critical, over budget" without repeating the numbers each plan was written with.

They go out once a day, after the daily snapshot, because the projection they carry only moves when the snapshot does. A shorter timer would repeat an unchanged number under a new timestamp, and evaluating on a page load would make the volume depend on who happened to be browsing.

The Event API key

Admin, New Relic holds a second key for this, beside the one everything else queries with. That is not a preference: the Event API takes an ingest key and refuses a user key, so the NRAK- key the rest of the application uses cannot send an event. Creating a new user key does not help, because it is the kind of key that is wrong, not that particular one. The field refuses an NRAK- outright rather than storing it and failing every night at two, and says so beside the field as well as at the Save button.

Get one in New Relic under API keys, Create a key, with key type Ingest - License or Ingest - Insert, in the account the events are going to. An insert key begins NRII-; a license key is a 40 character string ending NRAL. The account's existing license key works too, so there may already be one to use.

If a key that looks right is still refused with a 403, the usual cause is that it belongs to a different account from the events account ID below it.

Set the Events account ID to the account the events are written to, which is the account the ingest key belongs to and the one the alert condition is written in. The line under the box names that account back to you, so a number typed one digit short is obvious before anything is sent to it: a wrong account ID is refused by New Relic with the same 403 a wrong key gets, which sends you looking at the key.

All three buttons below the fields save what is typed in them first. A key typed above a button marked Send is meant to be the key that sends, so pressing the button stores it and then uses it, and the answer says so. Save settings still works the same way; you no longer have to remember it before testing.

Preview events builds what today's run would send without sending it, and Send now sends it immediately. The line under them says what the last scheduled run did, so a key that stopped working is visible on the screen that sets it rather than only in the log.

Test event sends one made-up event. It exists so the key, the account and the alert condition can all be proven before a real breach exists to prove them with, and so the answer to "is this wired up" does not have to be "wait for an account to go over".

The values are random but consistent with each other: a status and a side, a budget, a threshold, and a projection that really is past that threshold in that direction, with the daily rate and month to date derived from it. A chart built against a test event will not disagree with itself.

Two things make it safe to press. The account and plan names are invented, all beginning "NRAMP Test", so nobody reads one as an account of theirs going critical. And it carries test: true, which a real event never does, so a condition written as WHERE test IS NULL will ignore every test send. The screen says what it sent, and the full payload is in the browser console.

Pressing it does not touch the line reporting the last scheduled run, because a test overwriting that would hide a nightly send that has been failing all week.

Ingest and cost

Chargeback

Which budget owner and which cost center each New Relic account bills to. This is the table that turns ingest into an invoice somebody can reconcile.

Every column sorts, including Flags, which orders the table by how much the account check found and puts the worst rows at the top. Sorting changes the order on screen only: the stored table keeps its own order, so saving after a sort does not rewrite every row for the audit trail to report. A blank sinks to the bottom either way, and account ids sort as numbers rather than as text.

Editing

Edit puts every cell in play at once, because this is a sheet rather than a form. Removing a row is staged like any other change: nothing is written until you press Save, and Cancel discards the lot by reloading.

Every field is required: Account ID, Account Name, Subaccount Group, Budget Owner, GBL and Cost Center.

The table as a file

Export CSV, in the Archive, writes the table as it stands now, one column per field.

Import CSV replaces every row from a file, and keeps a copy of what it replaced first, noted New chargeback file imported by whoever did it and when. That copy is the way back, and it is taken before anything is written.

A file has to carry exactly the columns of the current schema, no more and no fewer, and is refused otherwise: the refusal names what is missing and what does not belong. A file missing a column would silently blank it for every account, and one carrying an extra column was written against a different schema, where guessing which of its columns to keep is how a chargeback table stops matching the bill. Column order does not matter, since they are matched by name.

What is in the rows is reported rather than refused, the same way the account check reports rather than rejects. The table legitimately holds rows a stricter rule would throw out, so an import says how many fields came in blank, how many rows have no account id and how many account ids repeat, and leaves the judgement to you. Exporting the table and importing it straight back leaves it exactly as it was.

Row ids are not taken from a file. An imported table is a new table, and reusing whatever numbers a spreadsheet happened to hold would tie new rows to the history of old ones they have nothing to do with.

Archive

Archive opens every kept copy of the table, newest first, with who kept it, how many records it holds and the note it carries.

Copies are kept three ways. One is taken automatically before every save, noted Saved before edit with the date and who saved. One is taken automatically before every restore, noted Saved before restore the same way, so restoring the wrong copy is itself undoable. And you can keep one by hand at any time: type a note saying why it is worth keeping and press Save a copy. A note is required, because a list of copies distinguished only by their timestamps is a list nobody can choose from.

Restore this copy stays disabled until a copy is chosen from the list, then asks you to confirm, quoting that copy's note, who took it and how many records it will replace. Restoring keeps the current table first, under its own automatic note.

The copies taken by hand are kept for good. The automatic ones are capped at thirty, so an afternoon of editing cannot bury the copies somebody meant to keep.

The account check

Every row is compared against the accounts your API key can see, automatically, each time the page opens and again after a save. The table appears first and the flags land a moment later, so a New Relic that is slow or unreachable costs the flags rather than the page; the line above the table says which happened.

Red means the row cannot be billed as written; yellow means it is worth fixing but the charge still lands.

  • no match, red. No New Relic account with that ID is visible to your key.
  • Account name mismatch. The ID matches an account, but New Relic calls it something else.
  • dash. The names match once dashes and spacing are normalized, which is the import artifact rather than a real disagreement.
  • duplicate account ID. The same account appears on more than one row, so two owners could be billed for it.
  • GBL missing or Budget Owner missing. The field is empty, so the charge has nowhere to land or nobody to receive it. A row missing both carries both chips. Only those two fields are flagged; a missing group is untidy but still billable.

It also reports how many live accounts have no row at all.

Nothing is rejected on the strength of a flag. A sub-allocation such as an Istio or GMAL line is deliberately not a New Relic account, and an account can be outside your key's scope while still being real. The check surfaces these; deciding is yours.

Creating an invoice

Pick a month, pick the accounts, give it a title, and generate. The month list offers only months with ingest recorded, and marks the current one as in progress, because billing a month that is still running would undercount it.

The invoice totals the month's ingest for those accounts at the rate set in Admin, and groups it by budget owner, since that is how it will be reconciled. It says so plainly when an account has no chargeback record, when a selected account reported nothing that month, and when no rate is set at all, rather than quietly leaving those out of the total.

Saving an invoice

A saved invoice keeps its own figures, including the rate it was priced at. Recomputing from live data whenever it was opened would let a later change to the rate, or to a budget owner, quietly restate a bill that has already gone out. Saved invoices can be reopened, printed and deleted; deleting one cannot be undone.

Audit trail

Every save, restore, add, delete and export is recorded with who did it and when, along with every invoice saved or deleted. It is append only, and the export is logged too, since a copy of this table leaving the app is itself worth knowing about.

Alerts and incidents

Alert Analysis

Pick any account and run a complete audit. Every policy and every NRQL condition is read to completion, with no sampling.

What the headline means

The percentage at the top is how many conditions can actually fire. Below 25 percent the report opens by saying the estate is effectively dark, because a policy that exists but is switched off produces no notification at all.

Sections

  • Coverage by policy group. Policies grouped by the prefix before a colon in their name, which is how these estates are organized in practice.
  • Empty policies. No conditions attached, so they cannot fire and cannot be told apart from working ones without opening them.
  • Policies without workflows. A workflow is what carries an incident to a destination. Without one, anything raised reaches nobody.
  • Conditions with bad NRQL. Queries that do not execute.
  • Disabled conditions, grouped by policy.
  • Conditions never fired. Enabled but silent. Silence can be correct for a rare failure, but a threshold that can never be met looks exactly the same.
  • Noisy conditions. High volume alerting that trains people to close without reading.
  • Development and test policies, incident preference, and recommendations generated from what the audit found.

Validating NRQL

The checkbox runs every condition query against the account to prove it parses. That is one query per condition, which is fine for a few hundred and slow for several thousand, so it is a choice rather than an assumption.

Creating tasks

Every row in sections 2 through 10 has a button that turns it into a task. Each section header has one that creates a single task covering every item in that section. Where a table is capped for display, the section task still covers the full set.

Results are cached for 7 days per account, and Run analysis opens that copy when there is one, which is why a large estate can appear at once. Re-run ignores the stored audit and queries New Relic now, which is the slow path and the one to use when the estate has changed. Admin can also have the cache filled for chosen accounts on a schedule, under Alerts analysis config.

Alerts and incidents

Incident Heatmap

A day of alerting across every account at once: which accounts raised incidents, in which hour, and what condition was behind them. Any one account can then be opened by condition, on the same hours.

The window

24 hours draws a column per hour. 7 days draws one per three hours, because a week of hourly columns is 168 of them, about five screens of dragging to read one row; at three hours a nightly batch still reads as a band. Columns start at midnight on your own clock either way, and the day is marked where it turns over. Each window is its own sweep with its own stored result, so going back to one already read is immediate, and an account left open is redrawn on the window you switch to.

What a cell counts

One row is an account, one column an hour on your own clock, oldest on the left, and the cell is the incidents that opened in that hour, or in those three on the week. Opened rather than open, so an incident that ran for nine hours is one cell rather than a row of nine. Colour is on a log scale: one account can raise more incidents in an hour than the rest of the estate raises all day, and a straight scale would leave every other account white.

Accounts that raised nothing are left out rather than shown as empty rows, and the summary says how many were quiet. The box at the top of the table filters the rows the same way the other search boxes do: it matches a name or an id, and several terms separated by commas show rows matching any of them. The condition map has its own, filtering conditions, and neither changes the colors, which stay keyed to the whole map so a filtered view can still be read against the legend. All, Critical and Warning switch which incidents the map, the totals and the color are counting.

One account by condition

Click an account name and the map is redrawn for that account alone, one row per alert condition, over the same hours. Alert Analysis carries an Incident Heatmap button beside its title that opens the same thing for the account being audited, so what could fire and what did sit one click apart. It answers the next question the account map raises: the account is loud, but is that one condition or fifty.

The fifty conditions raising the most get a row, and the tile says how many raised anything at all; the totals count the rows on screen. All accounts goes back. A cell here opens the incidents that one condition raised in that hour.

The condition's name opens the condition itself in New Relic, in a new tab. Rows are one per condition rather than one per name, because an account can run several conditions under the same name in different policies: this estate has two called "Memory on CSO Alert" raising 34,435 and 1,636 incidents a day. Where that happens each row shows the condition id beside its name.

Opening an hour

Click any cell that carries a number and it opens the breakdown of that account in that hour: the totals, every condition that raised something with its share of the hour, and the incidents themselves newest first, each with a link into New Relic.

A cell on the condition map opens the same thing for that one condition, and adds the entities it fired on. Entities are shown only there, because an hour of a whole account spreads across thousands of them: one measured hour held 7,068 incidents over 3,128 entities, where the twenty largest covered three per cent and read as a list of ties. Even for one condition the heading says what the chips cover, since New Relic leaves an incident with no entity out of the list entirely. A cell answers to the keyboard as a button as well.

The list holds the newest hundred and the conditions the fifty largest, which the headings say when either is reached. Nothing here is swept in advance: it is read when the cell is opened, because a day of cells nobody clicks would be over a thousand queries.

Where the figures come from

Every account the API key can see is asked for its own incidents. New Relic refuses a cross-account query naming more than 30 accounts, and more than five once it is faceted, and NrAiIncident carries no attribute saying which account raised it, so a single estate-wide query could neither run nor be split apart afterwards. Each account's query is a plain count over one event type, so a sweep of roughly 200 accounts finishes in seconds and costs almost nothing.

An account that cannot be read is listed at the bottom with the reason, rather than being counted as quiet.

Freshness

A sweep is kept for ten minutes and reused, because incidents are live data and the whole estate is re-read each time. The line above the summary says when the one on screen was taken; Re-run sweeps again now.

Conditions behind them

Under the map, the conditions raising the most incidents in the window, with the account each belongs to and its share of every incident counted. It is the fastest way to see that one noisy condition accounts for most of a day's alerting.

Dashboards, logs and metrics

Dashboard / Query Profiler

Pick an account to list its dashboards, then click one to see the queries behind its widgets, what they cost over the last day, and to run them. The CCU and time levels that color everything here are set in Admin, under CCU variance thresholds, and can be overridden per account.

The dashboard list

Ten rows at a time, every dashboard in the account, pages excluded. Search matches a dashboard's name or its owner, and several terms separated by commas show dashboards matching any of them. Dashboard, Owner, Popularity and Appeal each sort; the score columns start with the highest. Click any row to open its detail.

Two cards at the right of that row count the account's dashboards and how many nobody has run in the last 30 days. In Aurora Retail Prod that is 1,373 dashboards, of which about four in five are unused.

Popularity and appeal

Both come from the last 30 days of who actually ran each dashboard, and both count people rather than runs. A dashboard left open on a wall screen refreshes all day, so its run count says little about how many people rely on it.

Popularity is how many different people used a dashboard, from 1 to 10. It is on a log scale against the account's most used dashboard, because use is lopsided: most dashboards have one or two users and a few have dozens. A dashboard needs at least 10 users to score 10.

Appeal is whether people come back, from 1 to 10, however many there are. Half comes from the share of its users who used it on two or more separate days, and half from how many days its typical user used it, with full marks at eight days of 30. A dashboard one person keeps open every day can have high appeal and low popularity; read the two together.

A dashboard nobody ran in 30 days shows Unused. The Hide unused switch beside the search box leaves them out of the list; it is remembered, and it waits until usage has been scored. Scores are worked out once a day per account; the first visit of the day can take a little while and spends a few CCU. Hover a score for the numbers behind it.

Dashboard Detail

Cards across the top in three groups: the dashboard itself, its queries, pages and how many you have run here; the last 30 days, with its users and both scores; and the last 24 hours, with the runs New Relic recorded, what they cost in CCU and their average response time. Below them, one row per widget query, with the query text, what it cost over the last day, the result of running it here, and the button to run it.

New Relic, beside the heading, opens the dashboard itself in New Relic in a new tab.

Recorded, last 24h comes from New Relic's own records of the dashboard being used, matched to a widget by the text of its query. A run that a filter or a variable rewrote is recorded under different text and cannot be tied back to its widget, so a busy dashboard often shows a match for only some of its queries; the cards say how much of the day's cost the matched ones cover.

Run executes one query once, now, as written and over its own time window. It spends CCU, and can cost more or less than the dashboard's own runs. Its time and CCU are each marked OK, Warning or Critical against the account's levels, and the run's CCU multiplied by the dashboard's recorded runs gives an estimated cost per day, judged the same way. The row's left edge takes the worst of the three, and the table keeps its place while a query runs.

Narrowing the list, and running everything unmatched

Above the queries, the CCU and Execute filters narrow the list to queries at warning or critical level. A query takes the worst level among its figures, recorded and run now, so a critical query is not also counted as a warning. Choosing several filters shows queries matching any of them, and the filters stay set when you open another dashboard.

Run all, at the right of the filters, runs every query with no matching recorded run that can run as set, two at a time, after asking. Queries already run are skipped, and Stop starts no more while letting the runs in flight finish. Each run spends CCU.

Dashboard variables

Under a query that uses dashboard variables, each variable has a control that starts at the dashboard's default: a pull-down for a fixed list, a text box for a typed value, and for a list the dashboard fills from a query, a Load values button that runs that query once (it spends CCU) and fills the pull-down. Choose other values and Run uses them; Use defaults puts them back. Values are quoted and checked before they reach the query, and a fixed list accepts only its own choices. Recorded runs are matched only at the defaults, marked "at variable defaults". A multi-select variable set to every value cannot be expanded here, so choose the values to run it with.

A dashboard can narrow one list by another, listing markets only for the namespace chosen above them. Loading such a list uses what is chosen on that row, and where nothing is chosen, every value of the variable it depends on, which is what the dashboard itself does with "select all". Choosing a different value for the variable it depends on clears the narrowed list so it loads again.

Dashboards, logs and metrics

Log Partition Analysis

Which slices of a log partition would pay for a partition of their own, measured in the CCU a day they would save.

Why a partition, and not a WHERE clause

A query is charged for what it inspects, not what it matches, and what it inspects is the whole partition. Measured in Aurora Retail Prod over the same ten minutes, FROM Log inspected 76.9 million records for 0.0592 CCU, and FROM Log WHERE appName = 'gateway' inspected 76.3 million for 0.0591 while matching a tenth of them. Moving a slice into its own partition is therefore the only way to make the queries that do not want it cheaper, and they get faster in the same proportion.

What it measures

Pick an account and a partition and press Analyze. It reads the last 24 hours of queries New Relic recorded against that partition, works out what each one filters on, and takes a single ten minute sample of the partition to measure how much data each of those filters describes. It spends a fraction of a CCU and takes a few seconds; the result is kept for a day, and Re-run measures again.

Candidates are whole plans: split the partition along one attribute, a partition for each of its values, which is what people mean by partitioning logs. A partition each for orders, offers and payments, say, or one per market or per cluster.

Queries rarely name an attribute like capability, they name an application, so each value's applications are read from the data and a query naming one is routed to that value's partition. An application counts only when nearly all of its records carry the same value, since every application logs at every level and mapping on the last value seen would flatter a split by level enormously.

A value earns a partition by being large or by being heavily queried, so a small slice everyone asks for is not passed over: Aurora Retail Prod's menu capability is about 1% of the bytes and carries 38,000 CCU a day of queries. Values that differ only in capitals are treated as one, since a rule can name both spellings: Aurora Retail Prod logs its orders capability as both orders and Orders. Values that are neither stay in what remains, and the screen says how many.

Each candidate is ranked by the CCU a day it would save: the queries that can be routed to one partition, freed from scanning the others.

The biggest value is not where the saving comes from. market = 'us' is two thirds of that partition, so a query routed to it still reads two thirds of the data. The saving is in everything routed away from it.

Building a plan

Tick the splits you would write and press Measure this plan. Each rule is measured against what the rules above it have already taken, and every query is routed to the first partition that holds what it asked for.

When every rule tests one attribute for different values, the order makes no difference, because no record can match two of them; the screen says so and does not offer to reorder. Order matters when rules can overlap, such as one on appName beside one on capability: measured on Aurora Retail Prod, that pair scores 26% one way and 29% the other, because whichever rule applies first takes the shared records.

A measured plan also carries the rule New Relic itself needs for each partition: the matching clause, which is the rule's own NRQL with no FROM and no WHERE, and the NerdGraph mutation that creates it, with the account, the target partition name and a STANDARD retention policy already filled in. Field names come from the account's own schema rather than from documentation. NRAMP creates nothing: run the mutation with a key that may write log configuration, or type the clause into Manage data, Data partitions.

Queries that name nothing a rule could match read every partition whatever you do, and they are listed with what they cost. Narrowing the worst of them is often worth more than another partition.

What the change would break

Affected config entities, beside Measure this plan, lists every dashboard widget and alert condition reading this partition today and names the partition each would have to query instead.

Creating a partition rewrites nothing. A condition still reading the old partition keeps running and quietly stops seeing whatever moved, which is the real risk in this kind of change, so the list is worth reading before the rules go live.

Alert conditions are every condition in the account, read from the alerts API at no CCU cost. Dashboards are those New Relic recorded running in the last 24 hours, so a dashboard nobody opened is not listed. A query marked every partition filters on nothing a rule matches, so it cannot pick one and has to name them all.

Printing

Print partitions prints the plan: what each partition would hold, what it saves, and the rule to create it. Print everything adds the dashboards and conditions to rewrite, and the affected report has its own buttons for either list or both. Each is built as its own document rather than a screenshot of the screen, so the tables fit the page and nothing is cut off.

Before you act on it

Nothing here changes New Relic. Partition rules are created in New Relic's own console, they apply at ingest, so a plan changes tomorrow's queries and leaves yesterday's data where it is, and every query that wants the moved data has to name the new partition. That last point is the risk worth planning for: an alert condition still reading the old partition keeps working and quietly stops seeing what moved.

Dashboards, logs and metrics

Log Validator

Which of an account's logs anything actually asks for, and what the rest costs a month.

Every alert condition and every dashboard widget that reads a log table is collected, and a day of log ingest is measured stream by stream. A stream nothing names is one nobody has asked a question of: it is still paid for, still stored, and nothing on any screen would notice if it stopped arriving.

Choosing an account

The account list is the profiler's, largest by this month's ingest first, because log spend is concentrated in a handful of accounts. The screen opens on whatever was stored for that account by the nightly job. Validate now is an admin's button, and it asks before it starts. It reads every alert condition and every dashboard in the account and then measures every partition, which is thousands of calls to New Relic on the shared key: seconds on a small account, around six minutes on the largest. It is there for when an answer is needed sooner than tonight, and rarely for anything else, because the nightly job caches every account it covers and what is on screen is never more than 23 hours old.

What counts as required

A stream is required when at least one alert condition or dashboard widget names it, exactly or through a specific wildcard. Naming means a filter on the attribute that identifies the stream, not on what is inside it.

Most log queries filter on message text, and those read every stream in their table: counting them as asking for a stream would make everything required and the answer worth nothing. The screen says how many of an account's log queries are that kind. A wildcard that reduces to match-anything is ignored for the same reason. Two of them, entity.name LIKE '%' and service.name LIKE '%%', were enough to mark all 322 streams in Aurora Retail Prod required before they were excluded.

Widgets on the account's dashboards that read a different account are left out, since they say nothing about whether these logs are wanted.

Partitions, and why they change the answer

A query written FROM Log reads the default partition only. Where a rule routes an account's logs elsewhere, a query naming the app but reading the wrong table sees nothing, however precisely it names it. Each partition is therefore matched only against the queries that read that partition.

This is what the wrong partition verdict means, and it is not a saving. In Digital Production, 841 GB a day belongs to services named in log queries that read FROM Log while a rule sends their data to Log_GMAL: deliveryoptionselection-primary is named in 42 queries and none of them can see its EL region logs. Those are queries to fix. Dropping the data would make the blind spot permanent.

Keyed on says which attribute identified the streams in that partition. It is chosen per partition from what the records carry and what the account's own queries name, not assumed: a Lambda partition carries aws.logGroup and no appName, and keying the default partition on entity.name because more records carry it would report an estate as unused while every alert sits on appName.

A partition marked read wholesale is one queries read constantly without ever naming a stream in it, so nothing in it can be called unused and it is left out of the drop rules.

Two cards lead that table's title, each giving the records a month, the ingest a month and the money a month behind it.

The green one is the saving: the streams nothing mentions anywhere. It counts exactly what the drop rules cover, so the headline and the rules under it cannot disagree. The amber one is the question: streams named only in queries on other event types, where the service is watched through its metrics or its transactions and nobody reads its logs. Whether that is waste is a judgement about what someone would reach for after an incident, which is why it is not drawn in the color of money already found. Both leave out the partitions read wholesale.

Only untargeted, at the right of the same row, is on when the screen opens: the table lists the streams nothing asks for, largest first, each with the reason it is there. Turn it off for every stream the account ingests.

Note that turning it off changes little above the fold. Streams sort largest first and the large ones are nearly always required, so what the switch adds is a thousand rows below what is already on screen; the count beside the title is the quicker read.

The three kinds of untargeted

Named nowhere is the candidate to stop collecting: no condition and no widget mentions it at all.

Wrong partition is a monitoring gap, as above.

Non-log queries only means the name appears in queries on other event types, so the service is watched through its metrics or transactions while its logs are read by nobody. Whether that is waste is a judgement about what someone would reach for after an incident.

What to do about it

Three sections, one per kind of finding, each with a Create task button that puts everything in that section into one task: the rules in a section are one decision taken once, and whoever takes it wants all of it in front of them rather than three tasks to do one piece of work.

Suggested drop rules - Named nowhere is NRQL for the streams nothing mentions, chunked to stay inside New Relic's statement length and ready to paste in as a Drop data rule.

Suggested drop rules - Non-log queries is the same for the streams named only in queries on other event types. Written the same way, decided separately: something does watch those services, so judge each against what somebody would reach for after an incident.

Remediation - Wrong partition is not a drop rule and deliberately never becomes one. These queries name the stream correctly and read a partition that cannot hold it, so they return nothing and any alert built on them cannot fire. It gives the pair of partitions involved, the streams stranded between them, the conditions and dashboards doing it, and the FROM clause that fixes it. Dropping this data would make the blind spot permanent.

NRAMP creates nothing here, and a drop rule discards at ingest: what it drops is not recoverable.

The nightly job

Configured under Admin, Log validator config, and the accounts are the Log analysis column of the table under Cache scheduler. The Metrics Validator keeps its own job and its own column beside it. With that column left empty it validates the busiest 15 accounts by this month's ingest, which is where log spend actually is. It runs at 5am by default, an hour after the alert audit, because it reads every dashboard in every chosen account and has no reason to compete with the other three jobs for the same key.

Dashboards, logs and metrics

Metrics Validator

The Log Validator's question, asked of Metric: which of an account's metrics an alert condition or a dashboard widget actually asks for, and what the rest costs a month.

The screen is deliberately the same shape, because it is the same decision. What is underneath differs in two ways that matter.

How a metric query names what it wants

A log query names its stream in a WHERE clause. A metric query usually names its metric in the select itself: SELECT max(kafka_consumer_group_ConsumerLagMetrics_Value). Of the 20,071 metric queries in Aurora Retail Prod, only 827 carry a metricName predicate at all, so reading the WHERE clause alone would have reported almost every metric as unused.

Both are read, along with metricName wildcards, and the select is read through nesting: rate(sum(newrelic.timeslice.value), 1 minute) asks for the timeslice metric, not for sum.

Function arguments are collected loosely and matched strictly. uniqueCount(host.name) hands back an attribute rather than a metric, which costs nothing, because a candidate only counts once it matches a metric the account actually ingests. Being generous at the reading and strict at the match is what stops a spelling nobody anticipated from turning a watched metric into a drop candidate.

No partitions, and one key

metricName is on every record, so there is no key to choose, and metrics have no data partitions, so there is no wrong partition to be stranded in. FROM Metric is the whole of it.

That is why the third section here is about something else: metrics whose only reader is an alert condition that is switched off.

What the totals are

Totals are the sum of the metrics in the table rather than a separate measurement of the day. The two do not agree: bytecountestimate() sizes a faceted set differently from an unfaceted one, and on Aurora Retail Prod the day reads 10,005 GB whole against 8,667 GB summed across its 1,792 metrics. The point counts reconcile to within 35 of 2.9 trillion, so nothing is missing from the table; the screen says so where the difference is more than a couple of percent. Using the larger figure as the denominator would quietly understate how much of the account is required.

The three sections

Suggested drop rules - Named nowhere is NRQL for the metrics nothing asks for, by name or by wildcard.

Suggested drop rules - Other event types is for metrics whose names turn up only in queries on other event types. Check those queries before dropping: a name can be shared by a metric and an attribute somewhere else.

Remediation - Disabled conditions only is the metric side of the log screen's wrong partition, and the same kind of finding: a query exists and cannot serve anybody. These metrics are collected for an alert that is switched off. Turn the condition back on or delete it; no drop rule is offered, because dropping the metric settles that question by accident.

Each section carries a Create task button holding all of its recommendations.

Running it

Validate now is an admin's button and asks first, the same as the Log Validator and for the same reason.

The nightly job is its own, configured under Admin, Metrics validator config, with its accounts in the Metrics analysis column of the table under Cache scheduler. Left empty it takes the busiest 15 accounts by ingest. It runs at 6am, an hour after the log validator.

An account ticked in both that column and Log analysis has its conditions and dashboards read twice a night, once by each job. That is the cost of choosing the accounts for each separately, and it is a real cost on a large account: tick an account in one column only if the other answer is worth the read.

Planning the work

Account Assessment

One report for one account, assembled from every analysis this application knows how to do, written to be printed and sent to the person who pays for the account.

What it contains

The account's name and id, the date it was produced, and the budget owner, GBL and cost center from the Chargeback table.

An executive summary computed from the sections below it rather than written beside them, so the two cannot disagree. Where a section could not be measured, the summary says the number is unknown rather than counting a failure as a zero.

Ingest analysis for every data source in the account, against the same source four weeks earlier, and the remediation the variance screen would propose for each one, built from the same code so the two cannot drift apart.

Alert analysis: the full audit of every policy and condition, and beside it the incidents the account actually raised in the last 24 hours. What could fire and what did, in one place.

Dashboard analysis: how many dashboards were opened by nobody in the last 30 days, and a recommendation to review them. The dashboards themselves are on the profiler, which can sort and search them; a report is read once and the count is the finding.

Log analysis: how the account's logs are partitioned, with the candidate splits and what each would save, followed by the log validator's findings, its partitions and the drop rules for streams nothing mentions.

Metric analysis: the metric validator's findings and the drop rules for metrics nothing asks for.

What each section is worth

A section heading carries the saving acting on it would produce, out at the right margin. A section with no saving to offer leaves it off rather than printing zero: the alert audit is not about money, and a heading reading "Potential savings: $0" would say it had looked and found nothing.

Only what can actually be stopped is priced. That is not the same as everything nothing reads, and on some accounts it is nowhere near: a log partition queries read without ever naming a stream in it is left out of the drop rules, because nothing in it can be called unused. RTL Prod - US has 2.66 TB a day in one such partition against 5 GB the rules cover, so the report says $37 a month rather than $19,165. The wider figure is still stated beside it, as context rather than as a saving.

The partition savings are quoted in CCU rather than dollars, because a split changes what queries scan and not what is stored: the same logs are kept either way.

Producing one

Produce report runs every analysis in turn against that account. It is thousands of calls to New Relic on the shared key and takes minutes on a large account, so it asks first and says what it is about to do.

It reuses anything measured in the last day, which is what makes a second report quick, and Re-measure ignores all of that and runs everything again.

Running it is an admin's; reading a stored report needs only the module.

A section that fails is marked where its findings would be and the rest of the report still arrives. That is deliberate: a section quietly missing reads as "this account has no logs" rather than "nobody measured", and a report somebody forwards is the worst place for that distinction to go missing.

Print gives the document to send. The controls, the account picker and the status line are not part of it.

Planning the work

Workstreams

A workstream is a definition that holds one or more tasks. Use it to turn findings from the other screens into work someone owns.

All workstreams, or my tasks

The switch beside the filter changes what the page is about. All workstreams is the estate. My tasks is the work assigned to you, drawn from every workstream at once and shown in the same table you get by opening a workstream, so a task reads and edits the same either way.

Two tiles change meaning with the switch, deliberately. The priority counts become task priorities rather than workstream priorities, and Total workstreams becomes how many workstreams you have work in rather than how many exist. Tasks are matched on the Assignee field, against your username or your full name, so either spelling finds them.

The dashboard

Dashboard beside New workstream opens the same numbers drawn rather than tabulated: where the work sits, how it divides, and what all of it is worth.

Two pie charts, one for workstreams and one for tasks. Each can be cut three ways with the buttons above it: by status, by priority, and by who owns or is assigned it. Three rather than one because "distribution" means whichever of these you came to ask about, and all three are already counted.

It opens on status, unless everything is in the same status, which it is until people start moving work along. A pie with one slice reads as a broken chart rather than as an answer, so the page opens on a cut that divides and says underneath why it did.

Names are not a fixed list the way priorities and statuses are, so owner and assignee show the ten largest and gather the rest into one slice. Work with nobody's name on it counts as unassigned rather than being left out, which is usually the finding worth having.

Below the charts, a card per priority showing its tasks with the workstreams they sit in underneath, because four workstreams marked high can hold two tasks or two hundred. Then the two figures the work exists for: ingest saved a day, and value per year. Both say what they do not cover. A task added by hand carries no figures unless you type them, and a total that hides the tasks it skipped reads as the whole estate when it is a part of it.

Colors mean the same here as on the list page. High is red wherever it appears, complete is green, and names take a neutral ramp because a name has no meaning of its own to color.

In any task list, a task whose target date has arrived or passed and that is not complete shows its name in red with *followup under it. A completed task is never marked, however late it was. Clicking the marker opens a window to write a note into the task's work log and set a new target date; either on its own is enough, and the note is kept even if the date cannot be saved.

The Priority, Status, Impact and Value columns sort, in both a workstream's tasks and My tasks. Priority and status sort by what they mean rather than alphabetically, so the first click puts the most pressing first: high before low, open before complete. Impact and value start with the largest, and a task carrying no figure sinks to the bottom either way rather than crowding the top.

The summary

Totals across every workstream: how many there are, how they break down by priority, task progress, work log notes, and the savings and value rolled up from the tasks inside them. Realized counts only completed tasks, so projected value and banked value stay separate.

The Notes column counts work log entries across every task in that workstream, and the tile above totals them across all of them, so it is visible which work has been written up and which has only been recorded as a status.

Savings and value

Both are rolled up from the tasks in each workstream, using the figures each task captured when it was created from a finding. An asterisk marks a total that covers only some of the tasks, which happens when a task was created by hand and carries no figures. A partial sum is never presented as a complete one.

Deleting

Deleting a workstream deletes its tasks and every note on them. You are asked to confirm and told how many tasks go with it.

Click any row to open that workstream's tasks.

Planning the work

Workstream tasks

Every task in this workstream, with its own priority, status, assignee, impact and value. Percent complete is derived from the tasks rather than stored, so it cannot drift out of step with them.

Editing the workstream

The priority chip and the description are editable where they are read: click either one to change it. A description that is empty still offers itself, rather than leaving nothing to click.

Priority saves as soon as you pick one, and leaving without picking puts the chip back. The description has explicit Save and Cancel, since a paragraph is easy to lose by clicking away; Ctrl or Cmd with Enter saves, and Escape cancels.

The owner, in the line under the title, works the same way as priority. It offers the people who can sign in to NRAMP, anyone already named as an owner or assignee, and Someone else for a team or an outside contact. The list is the same one task assignees use.

Task detail

Click a row to expand it. You get the finding the task came from, including the account and source, the full detail captured at the time, editable status, priority, assignee and completion date, and a timestamped work log.

The work log

Notes are append only and stamped with the time they were added, so a task reads as a record of what happened rather than a single editable field. Ctrl or Cmd with Enter saves a note without reaching for the mouse.

Email template

The envelope button on a task drafts a note to whoever owns the budget for the account behind it, ready to paste into a mail client. It carries what the task is, what it costs per day and per year, the cost center it lands against, and the detail the finding was raised from.

The recipient comes from the chargeback table. Where there is nobody to name, the template says so and leaves a marked blank rather than inventing one: the task may have no account, the account may have no chargeback record or no owner set, or your own account may not have access to that table.

Printing

Print workstream produces a report of the workstream and every task in it, each expanded with its detail and work log. The print button on a task row produces that task alone.

Planning the work

Migration Tools

A workflow is a plan to copy New Relic configuration from one account into another. The table lists what is in flight and what is finished; Create workflow steps through defining one.

The four steps

  • Accounts. Name it, and pick the source and destination. Both boxes search as you type, across all your accounts by name or by account id, and confirm underneath what the text resolved to. An account cannot be both source and destination, and that is refused rather than left to fail halfway.
  • Entity Types. Which kinds of configuration to move.
  • Entities to Migrate. The specific entities, read live from the source account, grouped by kind. Each kind has its own search box and its own Only selected box, under its heading and scoped to it: searching the dashboards leaves the policies as they were. The boxes autocomplete from that kind's own names, the same way the account fields do, and commas separate terms so an entity matching any of them is shown. Only selected hides everything of that kind not already ticked, which is how you check what a long list adds up to. Each heading says how many of its kind are showing, and select all shown means what is on screen in that group rather than the whole kind.
  • Review. What it will do, then save it as ready.

Saving before you are finished

Save works from any step and keeps the workflow as a draft, however incomplete. Only a name is required, because a half-defined plan is worth keeping and demanding a destination before you could put it down would defeat the point. Clicking a workflow in the table reopens it at the step it was left on, and finishing updates that workflow rather than creating a second one.

Who can run one

Anyone with access to this module can define, save and review a workflow. Running a migration writes into a live account and is restricted to admins, enforced on the server rather than by hiding a button.

A run starts with a dry run, which reads every definition, rewrites what is account specific and checks it, creating nothing. Only then does the live run unlock, and it asks once more before writing. The confirmation names the state the dry run examined, so a workflow edited in between has to be previewed again rather than run on a stale preview. Progress is reported step by step and continues if you close the window. A workflow that has already run cannot be run a second time, and nothing here undoes what a live run created. A run where every item was created is complete. One that finished but could not create some items shows errors in orange, and can be retried: what was already created is skipped. Failed, in red, is kept for a run that stopped partway on an error of its own.

When either run finishes, Print summary lists what it migrated, or in a dry run what it would migrate: every item by type, its result, and any warnings. Saving it as a PDF names the file after the workflow. The list keeps a Print summary button on every workflow that has had a live run, so it can be printed again later; a dry run is not kept, so print that one before closing the window.

What can and cannot move

Only configuration: dashboards, alert policies and conditions, synthetic monitors and workloads. Hosts, APM applications, containers and Lambdas are not offered because they exist wherever their agent reports. Moving one means repointing the agent, not copying an entity, so listing them would offer something that cannot be done.

Entities come from the entity catalog rather than from recent activity, so a condition that has never fired is still found. A source account with thousands of one kind is listed only to a limit, and the wizard says so: selecting all selects what is shown, not the whole account.

Defining is not running

Creating a workflow saves the plan and marks it ready. Nothing is written to the destination account. Running one is deliberately a separate decision, and a separate piece of work, because copying into a live account is not something to trigger by finishing a form.

Deleting a workflow removes the plan only. It has no effect on either account.

Estate and people

User Audit

Enter an email address and press Audit. You get three things: what that person has been running, the groups they belong to, and the changes they have made.

Print puts the whole audit on paper, or into a PDF from the print dialog. The search box and the buttons come off, the report states who it is about and over what windows, and the groups are opened so the printed copy names every grant rather than counting them.

Query activity

New Relic logs every NRQL query it runs with the person behind it, which is how this can say what someone actually looks at rather than only what they are allowed to see. The panel reports where their queries came from, which parts of New Relic they were in, the dashboards they opened, whether any of it came through an API key, and the accounts they queried.

Two things the numbers do not mean. A dashboard re-runs its widget queries on every refresh, so these are query executions rather than visits: one dashboard left open on a wall screen outnumbers a person who opened ten. And the log is kept about 30 days, so this window is shorter than the audit records below however long a window you ask for.

Every account the person's groups grant is counted, and the busiest ten are broken down: someone can hold access to fifty accounts and run everything in one. A person who queried through a user key is shown the key's id and how much went through it, which is what separates a colleague reading dashboards from a script.

Group membership

Every group across every authentication domain, expanded to show what each one actually grants: account ID, account name, and role. Groups are collapsed by default because a senior user can sit in dozens of groups granting dozens of accounts each. The summary line gives the group count and the number of distinct accounts.

Organization wide roles

Groups carrying one are expanded automatically and flagged. Those govern the organization and user management rather than one account's data, and because they can grant account access the effective ceiling is the whole estate rather than the accounts listed.

What the audit records cover

Identity and access events are organization scoped and land in a single account, which is where these are read from. Logins, group changes, access grants, role updates and dashboard edits all appear. Actions taken inside another account are audited in that account and are not included here, so an empty list is not proof of inactivity.

Nothing on this screen is cached. Access changes are exactly what you would be checking for, so every lookup queries New Relic directly.

Estate and people

Environment As-built

An inventory of what is configured across every account your key can see. It collects on first open, stores the result, and only collects again once the stored copy is more than 7 days old, so opening this screen is normally instant. Re-collect ignores that window.

Configured, not busy

These are counts of things that exist, read from the entity catalog. That is deliberately different from counting what reported in the last 24 hours, which undercounts anything idle or disabled, and undercounts most severely where the problem is worst. One account here holds 344 alert conditions of which only a handful ever fire; an activity based count would report it as almost empty.

The arrows

Each summary tile compares against the previous stored inventory, with the change and percentage on hover. Direction is reported without judgment, since whether more hosts is good or bad depends entirely on what you expected. With only one inventory stored there is nothing to compare against and the arrows are absent.

What is excluded

Incidents and anomalies are left out of the entity totals. They are events rather than built things, and one account alone carries over three hundred thousand open issues, which would swamp every other figure.

Ingest per account comes from consumption data through a rollup account, so it costs one query for the whole estate rather than one per account.

Administration

Admin

Finding your way around

Admin is a menu on the left with one item per box: New Relic, Ingest variance thresholds, CCU variance thresholds, Cache scheduler, Cache status, Log analysis config, Alerts analysis config, My Links, Logs and Users. Choosing one shows that box on its own rather than scrolling through all of them, and the page remembers where you were so coming back lands in the same place. Save settings applies to whichever settings box is open; the Users box saves its own changes as you make them. Printing the page prints every section, since there is nothing to click on paper.

Alerts analysis config

Caches the alert audit for the chosen accounts, so Alert Analysis opens from the cache instead of auditing a whole estate while somebody waits. It runs weekly rather than nightly because an audit is kept for a week: running it nightly would pay the whole cost again to replace something still fresh. Validating every condition query is what the Alert Analysis screen asks for by default, and an audit that validated is stored apart from one that did not, so leaving that box ticked is what lets the screen open from the cache. It is also nearly all of the cost, at one live query per condition. Measured across the fifteen busiest accounts: 10,344 conditions in sixteen minutes, of which one account holding 7,668 of them took eleven.

Logs

Two logs, on one screen, behind the switch at the top: Application and Sign-ins. They are separate because they answer different questions, and mixed together a week of scheduled jobs buries the one sign-in worth looking at.

Every time on both screens is shown in your own zone with the zone named, so a row reads the same whoever opens it. The database holds UTC, which is what hovering a timestamp shows you.

### Application

What the application recorded for itself: each scheduled job starting and finishing, how much it cached, and anything that failed, with the failure detail kept in full. Filter by level, by which part of the app spoke, or by text. The click-through demo writes here too, under the source demo: one line each time somebody opens a part of it, naming the module and nothing about who was looking or which steps they read. The newest five thousand lines are kept and the rest fall away. The log is held in the database rather than the system journal, because the service is not permitted to read the journal back: it therefore reads the same wherever NRAMP runs, and survives a restart and a release.

### Sign-ins

Who came in, from where, and what was refused. One row per event, with the time, what happened, which account, the address it came from and the browser.

Five things are recorded. Signed in and signed out are the ordinary pair. Failed is a wrong password or a name matching no account, with which of the two underneath it. Refused is a correct password for an account that is disabled, which is worth telling apart from a wrong one: somebody still has working credentials. Locked out is recorded once, when repeated failures lock an address out, rather than for every attempt refused after it.

Passwords are never written, in any form, and a name that matched no account is cut short, since that box sometimes receives a password typed in the wrong place.

The address is whatever reached the application. In production nginx sits in front and passes the real client address through, and the line under the table says whether that is the case: without a trusted proxy every row would read as the proxy's own address, which is worse than none. The browser column shows the few words that identify it, with the full string on the hover.

The newest fifty thousand are kept, ten times the application log, because this is the one read backwards after the fact: a scheduled job that failed last month is history, a sign-in from an address nobody recognizes is not.

Clearing it is recorded in the application log rather than in this one, since a record of an erasure kept inside the thing erased is no record at all.

Sign-ins recorded before this screen existed are still in the application log under auth, where they used to go.

My Links

Shortcuts of your own, shown under My Links at the bottom of the launcher. They belong to your account rather than to the app, so everybody keeps their own set and nobody sees anyone else's. Only web addresses are accepted, and they open in a new tab.

Ticking the icon box fetches the site's logo from logo.dev. That service refuses every request without a key of its own, so the box stays off until one is set at the bottom of the My Links section. The key is the publishable kind, meant to sit in a web page, so it is not treated as a secret here; be aware that the browser showing a link sends that link's address to logo.dev to fetch the picture. A logo that will not load falls back to the first letter of the name rather than leaving a broken image.

API key

A New Relic user key, beginning NRAK. It is stored in NRAMP's database on the server and is never sent to any browser; the screen only ever shows whether one is set and its last four characters. One key serves everybody, so what NRAMP can see in New Relic is the same for every user; who may see it inside NRAMP is decided by the modules on their account. Verify key confirms it works and reports how many accounts it can see.

Budgets

The contracted allowances, for measuring consumption against. CCU is entered in millions; ingest defaults to petabytes with a selector for terabytes and gigabytes. Changing that unit converts the figure rather than reinterpreting it, so 1 PB becomes 1,000 TB and not 1 TB.

Both are stored in the units the data itself arrives in, absolute CCU and gigabytes, with the unit remembered only so the field reads back the way you typed it. Units here are decimal, a thousand to the step, because that is how New Relic bills. Leave either blank if there is no budget to measure against; blank is not the same as zero.

Ingest variance thresholds

The default applies to every account. Variance is measured as the current daily average against last month's daily average, and only growth counts, so a decline is always green. Trending starts at the fraction of the threshold you set here, which is where an account turns yellow.

CCU variance thresholds

The levels at which the Dashboard / Query Profiler marks a query yellow for warning or red for critical. There are three, judged separately: time per run, which is how slow a widget feels; CCU per run, the cost of one execution; and CCU per day, a query's cost multiplied by how often it runs, which is what actually drives the bill. Time is not a stand-in for cost. In Aurora Retail Prod, dashboard runs over five seconds carry only 31% of the CCU, while runs over 1 CCU carry 92% of it.

The defaults, 5 and 30 seconds, 1 and 10 CCU per run, and 100 and 1,000 CCU per day, were tested against a day of real dashboard queries in three large accounts, and chosen because they flag few runs that carry most of the cost. They do not suit every account: CDX Production's median dashboard query costs 6.25 CCU, so the default warning of 1 CCU per run would flag almost all of them. That is what per-account overrides are for. Any level an override leaves blank uses the default, and a warning can never be set above its critical.

An account whose dashboards query heavy data by nature can carry its own levels through an override, so the defaults stay meaningful everywhere else. How the profiler uses these levels, and what it colors, is described in that module's own help.

Ingest price

Optional. Set your effective rate per GB and the analysis screens state remediation value in dollars as well as gigabytes. Left blank, everything reports in gigabytes only and no rate is assumed.

Cache scheduler

An ingest analysis runs about a dozen queries against New Relic and takes long enough that people wait for it. This job runs those analyses overnight for the accounts you choose and stores the results, so the screens open from the cache instead. Results are kept for 24 hours, which is why the job repeats.

With no accounts ticked it takes the busiest ones by this month's ingest, in the number set beside the schedule. Inside each account it analyzes only the sources the analysis can explain and that have bytes in this month or last. The time is wall clock in the zone chosen, so 2am stays 2am through daylight saving, and the job runs two analyses at a time to stay light on a small instance.

Run now starts it immediately and it keeps going if you leave the page. After each run the section reports what it did: how many analyses, how long, how much processor time and memory it used, and the slowest source. It never runs while a snapshot refresh is collecting, since both want New Relic and the same processor.

Log analysis config

The same idea for the log record size screen, on its own timer, an hour after the analysis cache by default so the two do not compete for a small instance. For each chosen account it builds the size table for every log partition the account has, grouped the way that screen opens it, and stores the result.

Only the table is kept. The drill-down that shows example log lines is never cached, because those lines carry tokens and personal data that this app does not store. Asking for a different grouping, window or threshold also goes to New Relic live, since the stored copy is the one the job built.

Grouping a source's messages into patterns runs a regular expression over every message it logged, and the largest sources log far too much to read a whole day of: Digital Production's gateway records 3.6 TB in 24 hours, and asking for all of it fails with an error from New Relic rather than an answer. So the pattern table reads the most recent stretch of that source worth about 150 GB, and says which window it used; its record counts and shares are for that window. Everything else on the drill-down still covers the window you chose.

Cache every grouping, on by default, caches every way the screen can group every partition. Nothing is ever a cold query, and the cost is real: the two largest accounts carry most of it, and a grouping like hostname on a busy account takes about a minute to build. Unticked, the job caches the grouping each screen opens with and keeps warm any other grouping somebody has actually opened, which costs a fraction as much and leaves the rest to be built live on first view.

With it unticked, the job keeps warm any grouping somebody has opened. The first time you group a large account's logs by something like hostname it is a live query and can take a minute; after that the job refreshes it each night alongside the default, so it stays quick. Groupings nobody opens are never warmed, and one that stops being offered is dropped rather than queried forever. There is a ceiling per account, so exploring a big account for an afternoon cannot turn the nightly job into an hour of extra work.

The accounts table below the schedules carries a box for each job on every row, so an account can be in one, both or neither.

Which windows are cached

Only the 1 day window is filled overnight. The cache is keyed by account, source and window, so asking for 3 or 7 days never matches what the nightly job stored and the scan runs against New Relic every time. That is why the same source can open instantly at 1 day and take most of a minute at 7. The page says so when it happens.

Cache status

What the jobs have actually stored: how many entries of each kind, across how many accounts, how many are still within their life and how many have passed it, and how old the newest and oldest are. Ingest analyses and log record size tables are kept a day, alert audits a week. Something past its life is not lost; it is rebuilt the next time anyone opens it, or on the job's next run, whichever comes first. The figures are read when you open the section, so Refresh is there for watching a run fill the cache.

Daily snapshot hour

In the New Relic box. The hour, on the server's clock, after which the daily refresh runs; it runs once a day and only when an API key is stored.

API keys

Every user manages their own keys, from the name badge in the top bar rather than from here. A key calls the API as the person who made it and carries their access and no more, so losing a module narrows every key that person holds, at once. Disabling an account stops its keys the same way.

Users

Accounts for NRAMP itself, which are not the New Relic users the User Audit module inspects. Create one on the left, and the roster on the right shows each person's role, what they can reach, whether they have a live session, and when they last signed in.

A new account is created under an email address. Capitals do not matter when signing in, and a name already taken cannot be taken again in different capitals. Accounts made before that rule may carry a short handle instead, and those still sign in: a username cannot be changed, so refusing the format now would lock those people out rather than tidy anything up. To rename someone, create the new account and disable the old one, which keeps its history.

The password is typed twice and the two must match, said as you type rather than after the form is sent. The match is checked on the server as well, so it holds for anything that posts to the API directly.

An admin reaches every module and the whole of Admin. A superuser reaches every module, and inside Admin sees this page alone: every other section is listed, greyed and never opened for them. A superuser may create standard users and manage the standard users that exist, including granting any module to any of them. They cannot create or edit an admin or another superuser, cannot promote anyone, and cannot delete an account; disabling one is theirs to do. A standard user reaches only the modules checked for them, and never sees Admin. The checkboxes gray out for admins and superusers because the role already grants everything.

Access is enforced on the server, not just hidden in the page, so a standard user cannot reach a module by typing its address. The list shows twenty users at a time and scrolls beneath its heading for the rest. The search box above it narrows the list by username, name, role or module; separate several terms with commas to keep the users matching any of them. Click any row to change a username, a name, a role, adjust access, set a new password, or disable the account. A changed username has to be an email address nobody else is using, and is what they sign in with from then on. Disabling keeps the history and ends any live session; deleting removes the account outright.

The last active admin cannot be demoted, disabled or deleted, and you cannot remove your own admin access, since either would leave nobody able to manage users. Superusers do not count towards that, because they cannot reach the settings an admin can.

Welcome email. When an account is created, NRAMP can email the new user at their address: the NRAMP logo, your message, and a Sign in button that goes to the sign-in link. It never includes the password, which a member of the GTIO/Observability team sends separately in Microsoft Teams; the default message says so. Leave a blank line between paragraphs, and write {name}, {username} or {loginUrl} where the user's own details should appear.

Admins and superusers can edit the subject, the message and the sign-in link, and turn the email on or off. It is off until somebody turns it on, so accounts made before the mail server is set up are not emailed. The mail server is set by admins only, because its password is a company credential; it is saved and never shown again, so leave the field blank to keep it. Any SMTP service works: Amazon SES, Microsoft 365 or a relay. Use port 587 with STARTTLS or port 465 with TLS; port 25 is blocked from AWS by default.

Send me a test saves the form and sends the email to your own address, so you see exactly what a new user will get. If the email cannot be sent, the account is still created, and the message under Create user says why the email failed.

Rollup account

An account that reports consumption for the whole organization. It is found automatically on the first refresh, and setting it here skips that search.