Docs▸Getting started

Use an existing OpenTelemetry setup

Already instrumented? Point any OTLP exporter at Vigilon and keep every line of instrumentation you have. There is no Vigilon SDK or agent to install, and it can run beside your current backend. Two things take a few lines of code: recording errors that you catch, and reporting background jobs.

How it works

Vigilon ingests standard OTLP over HTTP. Anything that can emit OTLP, whether an OpenTelemetry SDK in any language or an OpenTelemetry Collector, can send to it directly.

Traces are all you need to send. Endpoint health, error groups, and job runs are all worked out from them. Metrics are accepted and optional: send them if you already have a metrics pipeline. Logs are not accepted.

SignalPathProtocol
Traces/v1/tracesOTLP/HTTP · protobuf
Metrics (optional)/v1/metricsOTLP/HTTP · protobuf
LogsNot accepted
Base URLhttps://ingest.vigilon.io
Auth headerAuthorization: Bearer [YOUR_API_KEY]

Point your SDK at Vigilon

The fastest path when nothing else receives your telemetry yet: standard OpenTelemetry environment variables, read by every OTel SDK. No config file changes.

Already exporting to another backend?

These variables hold a single destination, so setting them moves your telemetry to Vigilon and away from your current backend. To send to both, export from a Collector as shown below, or add a second OTLP/HTTP span exporter in your tracing setup, using the traces URL and auth header above.

shell
export OTEL_EXPORTER_OTLP_ENDPOINT="https://ingest.vigilon.io"
export OTEL_EXPORTER_OTLP_PROTOCOL="http/protobuf"
export OTEL_EXPORTER_OTLP_HEADERS="authorization=Bearer [YOUR_API_KEY]"
export OTEL_RESOURCE_ATTRIBUTES="service.name=checkout-api,deployment.environment=production"
Resource attributes

Vigilon organizes telemetry by service.name and deployment.environment. Set both so traces, endpoints, and errors group under the right service and environment.

If your service exports logs

OTEL_EXPORTER_OTLP_ENDPOINT applies to every signal your service exports, logs included, and Vigilon does not accept logs. If your service exports logs over OTLP, use the traces-only variables below in place of the first three above. Your logs keep going where they go today, and only traces are sent to Vigilon. The traces-only endpoint is used exactly as written, so it includes the full traces path.

shell
export OTEL_EXPORTER_OTLP_TRACES_ENDPOINT="https://ingest.vigilon.io/v1/traces"
export OTEL_EXPORTER_OTLP_TRACES_PROTOCOL="http/protobuf"
export OTEL_EXPORTER_OTLP_TRACES_HEADERS="authorization=Bearer [YOUR_API_KEY]"

Export from your Collector

Running an OpenTelemetry Collector? Add Vigilon as a second exporter and keep your current backend receiving everything it does today. Evaluate side by side and cut over whenever you're ready.

otel-collector.yaml
exporters:
  otlphttp/vigilon:
    endpoint: https://ingest.vigilon.io
    headers:
      authorization: Bearer ${env:VIGILON_API_KEY}

service:
  pipelines:
    traces:
      exporters: [otlphttp/vigilon, otlphttp/existing]

If your config has a metrics pipeline, you can list otlphttp/vigilon there too. Leave a logs pipeline as it is: Vigilon does not accept logs.

Record caught errors

Your framework instrumentation already records exceptions that escape a request. An exception that your code catches and turns into an error response is different: Vigilon sees a failed request with no error type, message, or stack trace, or no failure at all. To attach it, record the exception on the active span and set that span’s status to error. A small helper makes it one call:

from opentelemetry import trace
from opentelemetry.trace import Status, StatusCode


def record_error(error: BaseException) -> None:
    """Attach a caught exception to the current request or job run."""
    span = trace.get_current_span()
    if span.is_recording():
        span.record_exception(error)
        span.set_status(Status(StatusCode.ERROR, str(error) or type(error).__name__))

Call it inside the catch block, before you return the response. It does not throw or change the response; your code still controls that. It works the same way inside a job run, and outside a request or job it does nothing. In any other language, make the same two calls on the active span.

Report background jobs

Background jobs are not picked up from instrumentation alone. Vigilon reports one execution of a job, such as a cron tick, a queue message, or one loop iteration, as a job run when it receives a span like this:

SpanWhat Vigilon expects
Root spanNo parent, so each run is its own trace, even when the job starts inside a request.
KindINTERNAL
job.nameA name that stays the same across runs, such as sync-users. A different name per item, such as sync-user-42, creates a new job every time.
job.scheduleOptional. The cron expression your scheduler uses. It enables last-run and stale-job detection and does not run anything.
FailureStatus ERROR and the exception recorded on the root span itself, not on a child span.
Child spansCarry job.name too, and job.schedule when set. A long job’s root span is sent last, and this lets Vigilon recognise the run before it arrives.

This helper covers all of it. It starts the root span, records a failure and raises it again so your retry logic still works, and copies the job attributes onto child spans:

from contextlib import contextmanager
from typing import Optional

from opentelemetry import trace
from opentelemetry.context import Context
from opentelemetry.sdk.trace import SpanProcessor
from opentelemetry.trace import SpanKind

JOB_ATTRIBUTES = ("job.name", "job.schedule")


@contextmanager
def job_run(name: str, schedule: Optional[str] = None):
    """Report one execution of a background job as a Vigilon job run."""
    attributes = {"job.name": name}
    if schedule:
        attributes["job.schedule"] = schedule
    tracer = trace.get_tracer("vigilon.jobs")
    # context=Context() makes this the root of a new trace. An exception that
    # leaves the block is recorded on the span, marks it ERROR, and is re-raised.
    with tracer.start_as_current_span(
        f"job {name}", context=Context(), kind=SpanKind.INTERNAL, attributes=attributes
    ) as span:
        yield span


class JobAttributesSpanProcessor(SpanProcessor):
    """Copies the job attributes onto every span started inside a job run."""

    def on_start(self, span, parent_context=None):
        parent = trace.get_current_span(parent_context)
        parent_attributes = getattr(parent, "attributes", None) or {}
        for key in JOB_ATTRIBUTES:
            if key in parent_attributes:
                span.set_attribute(key, parent_attributes[key])

Wrap one execution of the job: with job_run("sync-users", schedule="0 3 * * *"): in Python, or await jobRun({ name: "sync-users", schedule: "0 3 * * *" }, () => syncUsers()) in JavaScript. Register the span processor once, where you set up tracing. In Python, call trace.get_tracer_provider().add_span_processor(JobAttributesSpanProcessor()). In JavaScript, add it to the spanProcessors list of your NodeSDK or tracer provider, ahead of the exporting processors. The zero-code register preload has no place for it. Short jobs are still reported, but a job that runs for more than a few seconds and makes database or outgoing calls may not be, so use a tracing setup file if you have jobs like that.

Scripts that exit

A script that ends right after its job can exit before the last batch is sent, and that batch holds the root span. Shut down or flush your tracer provider before the process exits.

Next steps