1 of 38

CONQUERING DISTRIBUTED CLOUD LOGGING WITHOUT BREAKING THE BANK�

your step-by-step guide

By Stathis Fylaktos, CTO & Founder @ DevWorx.tech

2 of 38

AGENDA

01

Examining the cost and usability of our solution compared to AWS Cloudwatch with/without Opensearch

02

Initial setup in AWS and configuring EFS to be used with Loki

03

Examining different approaches for persisting logs through Loki

04

Deploying a Loki instance, an OpenTelemetry collector instance, and an instance of Grafana to AWS using ECS and Fargate

05

Setting up a .Net API to write logs to our logging system

06

Short mention of the steps required to include tracing and metrics for a full observability suite

3 of 38

01

Examining the cost and usability of our solution compared to AWS Cloudwatch with/without Opensearch

4 of 38

01

​

Why are we using this opensource approach instead of relying on the native services for logging that AWS provides?

Two well-known approaches are CloudWatch and OpenSearch:

  • CloudWatch is relatively inexpensive for storage, but querying logs is limited and can get costly.
  • Logs are separated into log groups (one per app) and streams (one per instance), which makes searching cumbersome compared to systems like Loki or OpenSearch.

​

This is a screenshot of some event streams in CloudWatch :�

5 of 38

01

​

OpenSearch on the other hand is far more sophisticated approach. It is built using Elasticsearch and retrieving logs is fast while also supporting complicated queries for searching. It has a managed and a serverless option, with the first one being cheaper in the majority of cases and the second one being preferred for spiky workloads and simplicity.

​

We’ll use the cheaper, managed option for our comparison.

  • The main downside is that OpenSearch, for native AWS integration, requires CloudWatch subscription filters. These filters automatically stream logs to OpenSearch either via Lambda functions or Kinesis Data Firehose. While Kinesis Firehose is typically more cost-effective than Lambda for high volumes, both options add significant costs.
  • OpenSearch also requires EC2 instances for data and master nodes, plus EBS storage for indexes and replicas. These infrastructure costs are ongoing regardless of log volume.

​

  • The open-source logging pipeline with OTel, Loki, Grafana is also cloud-provider independent and provides full control over log retention and query performance.

​

  • It can also be enhanced with traces and metrics for a full observability suite.

​

  • Even without high volumes, when dealing with high log event frequency with small raw log size, OpenSearch's combination of subscription filter costs and infrastructure costs can easily be 3-4x more expensive than the OTel collector with Loki and Grafana solution running on ECS Fargate.

6 of 38

02

Initial setup in AWS and configuring EFS to be used with Loki

7 of 38

02

​

We’ll create an EFS file system.

​

We’ll use its ID in the EFS Policy below:

8 of 38

02

​

​

Now we’ll create the role to handle the task execution in ECS Fargate with the following permissions:

​

9 of 38

02

​

​

This is the CustomEFSPolicy we will need to create:

10 of 38

02

​

​

We’ll now create the ec2InstanceRole with permissions to mount the specified EFS file system and to describe its mount targets (CustomEFSPolicy).

11 of 38

02

​

​

​

Next, we’ll create the EC2. Some key points to have in mind:

  • Create key pair certificate during EC2 instance creation to access the instance later.
  • Also attach the ec2InstanceRole you created earlier.
  • We’ll use the same security group for all our services in the VPC for simplicity.

​

​

​

Open your terminal, go to the directory where your certificate was saved and run chmod 400 <certificate-name>.pem for read permissions.

​

​

​

We need to ssh into the EC2 instance using the the key pair we created

ssh -i <certificate-name>.pem ec2-user@your-EC2-instance-public-ip

​

Add SSH port 22, 0.0.0.0/0 (testing) to the inbound rules of the security group, along with another entry in the inbound rules to allow all our services to communicate with each other in any port, using the default security group as source:

12 of 38

02

​

​

​

Once connected to the EC2 instance, run the following commands to mount the EFS to /mnt/efs and create the folders needed by Loki

​

sudo mount -t nfs4 -o nfsvers=4.1 fs-066b2ba3379abea8d.efs.eu-south-1.amazonaws.com:/ /mnt/efs

​

sudo mkdir -p /mnt/efs/loki/rules

sudo mkdir -p /mnt/efs/loki/chunks

sudo mkdir -p /mnt/efs/loki/wal

sudo mkdir -p /mnt/efs/loki/index

sudo mkdir -p /mnt/efs/loki/boltdb-cache (tsdb-cache when using TSDB)

sudo mkdir -p /mnt/efs/loki/compactor

sudo chown -R 10001:10001 /mnt/efs/loki

sudo chmod -R 755 /mnt/efs/loki

​

​

​

This ensures the EFS is mounted: df -h | grep efs

13 of 38

02

​

We need to create the private repositories to ECR so that our images are pushed there, just like in screenshot below:

14 of 38

03

Examining different approaches for persisting logs through Loki

15 of 38

03

​

BoltDB-shipper with EFS (v11 storage schema)

Stores indexes, chunks and WAL on a shared filesystem (e.g. EFS)

​

​

​

In this mode :

​

​

​

​

​

​

​

​

​

Pros:

  • Simple to set up.
  • No dependency on S3.
  • Minimal moving parts (no shipping/compaction process).

​

Cons:

  • Poor scalability — EFS latency makes queries slower as data grows.
  • Expensive if all logs stay on EFS beyond a few GBs/month, especially with high query rates that drive up throughput costs.
  • No structured metadata support.

​

When to use:

Suitable for development, testing, or small-scale, low-query-volume deployments where simplicity is important.

INGESTER

EFS

Logs to WAL for durability

Indexes (BoltDB files) and chunks to EFS

There is no shipping of index files or chunks to object storage.

16 of 38

03

​

BoltDB-shipper with EFS and S3 (v11 storage schema)

Uses an external store (e.g. S3) for chunks and indexes and EFS for WAL.

​

In this mode:

​

​

​

​

​

​

​

​

​

​

​

​

​

Pros:

  • Works reliably and is cheaper than using just EFS without S3 for small to medium workloads.
  • Durable storage (chunks and indexes in S3, WAL safety net on EFS).

​

Cons:

  • Queries get slower at scale (large index files).
  • More complex than EFS-only mode.
  • No structured metadata support.

​

When to use:

Best for small to medium production workloads (tens to hundreds of GB/month) where you want durability in S3 but can still tolerate slower queries at scale.

INGESTER

EFS

Logs to WAL for durability

S3

Logs are buffered in memory and grouped into chunks.

Indexes (BoltDB files) are first written to EFS

Indexes and chunks periodically shipped from EFS to S3 by the BoltDB-shipper.

After chunks and indexes are safely persisted, the corresponding WAL entries are truncated.

Chunks flush → written to EFS

17 of 38

03

​

TSDB with EFS and S3 (v13 storage schema)

Stores chunks, indexes, and metadata together in TSDB blocks.

Supports structured metadata and native OTLP ingestion.

​

In this mode:

​

​

​

​

​

​

​

​

​

​

​

​

​

Pros:

  • Scalable and more efficient storage layout (blocks).
  • Faster queries, especially at larger volumes (even faster if compactor is setup to help with querying)
  • Supports structured metadata.
  • Actively developed. All new features target TSDB.

​

Cons:

  • Operationally heavier compaction process (more vCPU/memory, more S3 traffic).

​

When to use:

Best for medium to large production workloads (hundreds of GBs to multi-TBs/month), where query performance, structured metadata, and long-term scalability matter.

INGESTER

EFS

Logs to WAL for durability

S3

Logs (chunks) are buffered in memory.

TSDB shipper uploads blocks from EFS

Loki fetches blocks from S3. It can cache them on EFS, so repeated queries are faster

Local staging files can be deleted once upload succeeds.

Chunks flushed → written as blocks to EFS

18 of 38

04

Deploying a Loki instance, an OpenTelemetry collector instance, and an instance of Grafana to AWS using ECS and Fargate

19 of 38

04

​

For Loki we need the following Dockerfile:

�# Use the official Loki image as the base

FROM grafana/loki:3.4.5

​

# Copy the Loki configuration file into the container

COPY loki-config.yaml /etc/loki/loki-config.yaml

​

# Define Loki's startup command

CMD ["-config.file=/etc/loki/loki-config.yaml"]

​

​

Note:

  • In case we use any of the storage options for Loki that include S3, we’ll first need to create a bucket preferably with ”Block public access”. The bucket in our example is “test-files-persist-demo”

​

20 of 38

04

​

And a file called loki-config.yaml that contains the Loki configuration and is the following:

BoltDB-shipper with EFS (v11 storage schema)

​

BoltDB-shipper with EFS and S3 (v11 storage schema)

​

TSDB with EFS and S3 (v13 storage schema)

​

21 of 38

04

​

These are the steps to upload the docker image to ECR:

​

Example commands for MacOS:

​

Install aws cli:

  • curl "https://awscli.amazonaws.com/awscli-exe-linux-x86_64.zip" -o "awscliv2.zip" unzip awscliv2.zip sudo ./aws/install

​

  • aws configure list (Check aws cli settings)

​

  • export AWS_ACCESS_KEY_ID=…
  • export AWS_SECRET_ACCESS_KEY=…
  • export AWS_REGION=…

​

  • Add AmazonEC2ContainerRegistryPowerUser policy to the user used to upload images to ECR and AmazonECS_FullAccess to work with ECS clusters and services

​

Login to AWS from the command line:

  • aws ecr get-login-password --region eu-south-1 | docker login --username AWS --password-stdin <registry-uri>

(registry uri example : 897729145798.dkr.ecr.eu-south-1.amazonaws.com)

​

​

We navigate to the folder that the Dockerfile exists and run:

  • docker buildx build --platform linux/amd64 -t <private-repository-Loki-uri>:v1 --push .

​

​

22 of 38

Once its uploaded, we create a task definition for the Loki service:

04

​

23 of 38

04

​

Now that the task definition is in place, we’ll create a Loki service, by using the task definition and by enabling the service discovery.

​

​

​

For example, if we use “loki-serv-discovery-test” as a service discovery name and “namespace-test” as a namespace we can then use http://loki-serv-discovery-test.namespace-test:3100 to access Loki from another service in the VPC (OTel collector, Grafana)

​

​

​

This is done to have a DNS name that the services can use, instead of complicated and costly solutions like assigning a load balancer and a target group to the service and then a DNS name to the load balancer.

​

24 of 38

04

​

We’ll continue the process for the Otel collector. This is the Dockerfile we’ll use:

# Start with the official OpenTelemetry Collector Contrib base image

FROM otel/opentelemetry-collector-contrib:0.121.0

​

# Copy custom configuration file into the container

COPY ./otel-collector/otel-collector-config.yml /etc/otel/config.yaml

​

# Expose the necessary ports

EXPOSE 8889 13133 55679 4317

​

# Set the default command to load the custom configuration

CMD ["--config", "/etc/otel/config.yaml"]

25 of 38

04

​

And this is an example configuration for the OTel collector named otel-collector-config.yml:

  • For the endpoint we use the DNS name from Loki’s service discovery

Native OTLP ingestion

(TSDB schema v13)

Loki native (non-OTLP) ingestion (BoltDB-shipper schema v11)

26 of 38

​

Once its uploaded, we create a task definition for the OTel service (screenshot on the right):

​

04

​

Now, if we are still logged in to AWS, we can try to push the image directly to ECR:

​

  • docker buildx build --platform linux/amd64 -t <private-repository-OTel-uri>:v1 --push .

Similarly to Loki, now that the task definition is in place, we’ll create an OTel collector service, by using the task definition and by enabling the service discovery.

27 of 38

04

​

Now we’ll push the Grafana image to AWS ECR. This is the Dockerfile:

​

FROM grafana/grafana:12.1.0

​

# Copy the custom datasource configuration file into the container

COPY grafana-datasources.yml /etc/grafana/provisioning/datasources/datasources.yml

​

# Expose Grafana's default port

EXPOSE 3000

​

# Optionally set environment variables to enable anonymous access:

# You can also define this in your ECS task definition.

ENV GF_AUTH_ANONYMOUS_ENABLED=”false"

ENV GF_AUTH_ANONYMOUS_ORG_ROLE="Admin"

​

# Use the default command for Grafana

CMD ["grafana-server"]

28 of 38

04

​

And this is the grafana-datasources.yml file needed by Grafana:�

​

Again, if we are still logged in to AWS, we can try to push the image directly to ECR:

​

  • docker buildx build --platform linux/amd64 -t <private-repository-Grafana-uri>:v1 --push .

29 of 38

04

​

The task definition for the Grafana service:

Similarly to Loki and the OTel Collector, we’ll create the Grafana service, by using the task definition and by enabling the service discovery.

30 of 38

04

​

We’ll now go to our security group and add the entry for port 3000. This is what you should see now in the security group’s inbound rules:

Please bear in mind that the settings used throughout this course are for development purposes. In a production env, Grafana should run in a subdomain with an SSL certificate in place, Loki should run in different security group allowing traffic from OTel collector and Grafana, etc.

31 of 38

05

Setting up a .Net API to write logs to our logging system

32 of 38

05

​

Simple example of setting up our .Net API to write logs to our logging system through the Serilog OpenTelemetry sink.

​

This is how our .csproj looks like:

�

​

​

​

​

​

​

​

​

​

​

​

​

​

​

33 of 38

05

​

And this is how Program.cs looks so that Serilog is setup as the logging provider, while configuring the OpenTelemetry sink with OtlpProtocol.Grpc so that our logs are being sent to the Otel collector through GRPC:

34 of 38

05

​

This is a screenshot from Grafana displaying test logs (navigate to the Fargate task’s public IP at port 3000):

35 of 38

06

Short mention of the steps required to include tracing and metrics for a full observability suite

36 of 38

06

​

To build a complete observability suite logs are not enough. We need traces and metrics.

​

Why do we need traces and how do they help?

​

We want to correlate logs under the same transaction.

​

  • When using logs, even with structured logging, in distributed scenarios each service may write logs in a different way.

​

  • Traces include spans (which have a duration, are not points in time like logs) and the structure of the span will remain the same across services.

​

  • In a trace (which is a process such as the creation of a user), each span shows an operation, includes hierarchy (what is described in the span happened because another thing happened), and belongs to a trace. This way we have the correlation of everything that belongs to the trace.

​

Let’s think of what is Instrumentation: adding code or libraries to your application so it actually produces telemetry data (logs, traces, metrics) in a standard format (e.g., OTLP) that your observability stack can consume.

​

If you add the right packages and configure OTel in your .NET app, many spans are recorded automatically without changing your business code.

37 of 38

06

​

Why do we need metrics?�Metrics are numeric measurements that can be aggregated over time (e.g., counters, gauges, histograms).�They let us answer questions like:

  • What was the average response time of an endpoint?
  • How many errors occurred in the last day for this API?

​

Using OpenTelemetry for metrics ensures each measurement is enriched with resource attributes (service name, environment, version), making them easier to correlate with logs and traces.

�With the code below in Program.cs and two instrumentation packages (OpenTelemetry.Instrumentation.AspNetCore, OpenTelemetry.Instrumentation.Http), we can start exporting traces and HTTP-related metrics to the OpenTelemetry Collector:

​

​

​

​

​

​

​

​

​

​

​

​

​

​

​

​

From then on, we can use Jaeger or Tempo for traces and Prometheus for metrics. We will need though to make additions to the OTel collector configuration and task definition and also to Grafana, in order to have access to this information. The Jaeger traces are also visible from Jaeger’s own UI.

38 of 38

THANK YOU

  • By Stathis Fylaktos, CTO & Founder @ DevWorx.tech

​

​

​