DEV Community

Tony Goodchild
Tony Goodchild

Posted on Originally published at latticegrid.dev

Put Splunk Search Results in Your Own App's Data Grid

Products that run on infrastructure have a strong chance of having their events in Splunk. The people who need to see those events are often in your application, not in Splunk, and often without a Splunk seat. The usual answers are an iframe around a Splunk dashboard, or a hand-built table over the REST API that re-implements filtering and paging badly.

Lattice Grid has a Splunk adapter that does the translation for you. This post shows what it does, what it deliberately does not do, and the three ways to handle the credential depending on who the page is for.

The shape of it

Splunk's on-disk index format is closed. There is no file to read and no bucket to open. The supported surface is the search REST API, so that is what the adapter speaks: it dispatches a search job, watches it, reads a window of results, and cancels the job the moment the grid asks for something else.

import { createGrid, createPushdownSource, splunkAdapter } from '@toclocoinc/lattice-grid';
import * as compute from '@toclocoinc/lattice-grid';

const grid = createGrid(el, {
  columns,
  rowKey: 'event_id',
  source: createPushdownSource({
    adapter: splunkAdapter({
      url: 'https://splunk.example.com:8089',   // the management endpoint, not Splunk Web
      search: 'index=security sourcetype=firewall',
      earliest: '-24h',
      headers: { Authorization: 'Bearer ' + token },
    }),
    compute,
    pageSize: 100,
  }),
});
Enter fullscreen mode Exit fullscreen mode

Give it a base search and the grid is live over it. Typing into the filter row, clicking a column header to sort, or scrolling to the next page each becomes a Splunk search job. The browser only ever receives the rows in view.

What a filter turns into

SPL has two ways to say "keep the events where this holds". | where is an eval expression: exact, case-sensitive, applied after the events have been read. The search command's own terms are matched against the index, case-insensitively, with wildcards. The adapter emits the search form, for two reasons. It is the faster one, because Splunk seeks rather than reads and discards. And it is the one whose result matches what the same filter means everywhere else in the grid, whose filter kernel compares strings case-insensitively unless a condition asks otherwise.

That choice has a cost, and the adapter names it rather than hiding it. A condition that asks for caseSensitive cannot be honoured by a case-insensitive matcher, so it is pushed case-insensitively and the grid says so once in the console. Returning quietly wider rows is the failure this adapter exists not to have.

Values never leave their quotes. Strings are wrapped and escaped, including the asterisk, because an unescaped * inside a quoted SPL value is a wildcard. A user who types foo*bar into a filter box gets events containing foo*bar, not everything from foo to bar. The wildcards that contains, startsWith and endsWith need are added outside the escaped text, so the adapter's wildcards are wildcards and the user's are literal. Field names are validated against a pattern and refused by name otherwise, since the identifier is the one part of an SPL expression that cannot be quoted away.

Time filters become the job's window

Splunk's earliest_time is inclusive and latest_time is exclusive. So a >= on the time column maps onto earliest_time exactly and a < maps onto latest_time exactly. The adapter lifts those two out of the filter expression and makes them the job's own window, which is the difference between Splunk seeking to a time range and Splunk reading everything and discarding most of it.

The bounds that do not map exactly, > and <=, and any time condition inside an or, stay in the expression as a numeric comparison on the time field. That is exact, just not as fast. Correct first, fast where correctness allows.

Seeing what was pushed

After each query the source reports what ran in Splunk and what, if anything, stayed in the browser:

grid.on('rows:changed', () => {
  const plan = source.lastPlan();
  console.log(plan.unpushed.length ? 'pushed part' : 'pushed all', plan);
});
Enter fullscreen mode Exit fullscreen mode

For most filter shapes the answer is "pushed all". When it is not, the plan tells you which condition was finished client-side and why, so there is never a silent gap between the filter the user set and the rows they see.

Where the credential lives

The adapter holds no credential of its own. Every request is made with the fetch and headers you supply. It stores nothing, prompts for nothing and refreshes nothing, because an adapter that cached a token would be a place for it to leak from. Three shapes cover the three kinds of page.

Production: a proxy in front of Splunk. Your application authenticates the visitor its own way. A small backend holds the Splunk service token, forwards only the search endpoints, restricts index= to what that visitor may see, rate-limits, and adds the cross-origin headers for your origin. The grid points at the proxy and never sees a token:

const adapter = splunkAdapter({
  url: 'https://app.example.com/splunk',   // your backend, not Splunk
  search: 'index=security sourcetype=firewall',
  earliest: '-24h',
});
Enter fullscreen mode Exit fullscreen mode

The proxy is about forty lines of Node, or a Lambda behind CloudFront. It forwards POST /services/search/v2/jobs, the job's own GET, /results, /export and the DELETE that cancels a job, and refuses everything else. The reference implementation is in the guide linked at the end.

Internal tools: a per-user session key. Behind your own login, each user signs into Splunk with their own account, and the page holds their short-lived session key in memory:

const key = await fetch(url + '/services/auth/login', {
  method: 'POST',
  body: 'username=' + user + '&password=' + pass + '&output_mode=json',
}).then((r) => r.json()).then((j) => j.sessionKey);

const adapter = splunkAdapter({
  url,
  search: 'index=security sourcetype=firewall',
  headers: { Authorization: 'Splunk ' + key },
});
Enter fullscreen mode Exit fullscreen mode

What leaks if the page is compromised is one visitor's own access, and it expires on its own.

Trying it out: your token in the page. Fine for a developer pointed at a development instance, which is what the first example on this page is. Never ship a shared token to visitors. A token grants everything its role can do, across every index it can see, to anyone who reads the page.

A token that expires goes in a fetch wrapper rather than in headers, since the wrapper is the only thing that can mint a fresh value per request.

When you only have an export

Without REST access to the instance, the grid still works from exported results, and exporting is something Splunk does well. From the UI, the CLI, or the REST results endpoint with output_mode=csv or json:

splunk search "search index=security sourcetype=firewall earliest=-24h" \
  -output csv -maxout 0 > firewall-24h.csv
Enter fullscreen mode Exit fullscreen mode

A CSV goes through grid.import, which infers each column's type and previews the mapping before anything lands. A JSON or NDJSON export goes through createUrlSource, which streams NDJSON line by line and renders the first rows before the rest has arrived. The REST route is the one worth automating: a scheduled job that lands a fresh export, with nobody exporting anything by hand.

There is a third route for teams already using Splunk's Ingest Actions to land events in S3 as NDJSON or JSON. What arrives there is ordinary files in a bucket you own, and the grid's DuckDB adapter reads them directly, with filters pushed into DuckDB the same way they are pushed into Splunk here. That works because Ingest Actions writes an open format to a destination you control. Pointing the grid at Splunk's own storage is not something it does.

Try it

There is no public Splunk instance or credential a website can publish, so the live demo runs the adapter against a mock of the search REST API. Everything above the mock is the real code path: https://www.latticegrid.dev/demos/pushdown-splunk/

The full guide, including the reference proxy: https://www.latticegrid.dev/docs/sources/splunk/

Top comments (1)