CONQUERING DISTRIBUTED CLOUD LOGGING WITHOUT BREAKING THE BANK�
your step-by-step guide
By Stathis Fylaktos, CTO & Founder @ DevWorx.tech
AGENDA
01 | Examining the cost and usability of our solution compared to AWS Cloudwatch with/without Opensearch |
02 | Initial setup in AWS and configuring EFS to be used with Loki |
03 | Examining different approaches for persisting logs through Loki |
04 | Deploying a Loki instance, an OpenTelemetry collector instance, and an instance of Grafana to AWS using ECS and Fargate |
05 | Setting up a .Net API to write logs to our logging system |
06 | Short mention of the steps required to include tracing and metrics for a full observability suite |
01
Examining the cost and usability of our solution compared to AWS Cloudwatch with/without Opensearch
01
Why are we using this opensource approach instead of relying on the native services for logging that AWS provides?
Two well-known approaches are CloudWatch and OpenSearch:
This is a screenshot of some event streams in CloudWatch :�
01
OpenSearch on the other hand is far more sophisticated approach. It is built using Elasticsearch and retrieving logs is fast while also supporting complicated queries for searching. It has a managed and a serverless option, with the first one being cheaper in the majority of cases and the second one being preferred for spiky workloads and simplicity.
We’ll use the cheaper, managed option for our comparison.
02
Initial setup in AWS and configuring EFS to be used with Loki
02
We’ll create an EFS file system.
We’ll use its ID in the EFS Policy below:
02
Now we’ll create the role to handle the task execution in ECS Fargate with the following permissions:
02
This is the CustomEFSPolicy we will need to create:
02
We’ll now create the ec2InstanceRole with permissions to mount the specified EFS file system and to describe its mount targets (CustomEFSPolicy).
02
Next, we’ll create the EC2. Some key points to have in mind:
Open your terminal, go to the directory where your certificate was saved and run chmod 400 <certificate-name>.pem for read permissions.
We need to ssh into the EC2 instance using the the key pair we created
ssh -i <certificate-name>.pem ec2-user@your-EC2-instance-public-ip
Add SSH port 22, 0.0.0.0/0 (testing) to the inbound rules of the security group, along with another entry in the inbound rules to allow all our services to communicate with each other in any port, using the default security group as source:
02
Once connected to the EC2 instance, run the following commands to mount the EFS to /mnt/efs and create the folders needed by Loki
sudo mount -t nfs4 -o nfsvers=4.1 fs-066b2ba3379abea8d.efs.eu-south-1.amazonaws.com:/ /mnt/efs
sudo mkdir -p /mnt/efs/loki/rules
sudo mkdir -p /mnt/efs/loki/chunks
sudo mkdir -p /mnt/efs/loki/wal
sudo mkdir -p /mnt/efs/loki/index
sudo mkdir -p /mnt/efs/loki/boltdb-cache (tsdb-cache when using TSDB)
sudo mkdir -p /mnt/efs/loki/compactor
sudo chown -R 10001:10001 /mnt/efs/loki
sudo chmod -R 755 /mnt/efs/loki
This ensures the EFS is mounted: df -h | grep efs
02
We need to create the private repositories to ECR so that our images are pushed there, just like in screenshot below:
03
Examining different approaches for persisting logs through Loki
03
BoltDB-shipper with EFS (v11 storage schema)
Stores indexes, chunks and WAL on a shared filesystem (e.g. EFS)
In this mode :
Pros:
Cons:
When to use:
Suitable for development, testing, or small-scale, low-query-volume deployments where simplicity is important.
INGESTER
EFS
Logs to WAL for durability
Indexes (BoltDB files) and chunks to EFS
There is no shipping of index files or chunks to object storage.
03
BoltDB-shipper with EFS and S3 (v11 storage schema)
Uses an external store (e.g. S3) for chunks and indexes and EFS for WAL.
In this mode:
Pros:
Cons:
When to use:
Best for small to medium production workloads (tens to hundreds of GB/month) where you want durability in S3 but can still tolerate slower queries at scale.
INGESTER
EFS
Logs to WAL for durability
S3
Logs are buffered in memory and grouped into chunks.
Indexes (BoltDB files) are first written to EFS
Indexes and chunks periodically shipped from EFS to S3 by the BoltDB-shipper.
After chunks and indexes are safely persisted, the corresponding WAL entries are truncated.
Chunks flush → written to EFS
03
TSDB with EFS and S3 (v13 storage schema)
Stores chunks, indexes, and metadata together in TSDB blocks.
Supports structured metadata and native OTLP ingestion.
In this mode:
Pros:
Cons:
When to use:
Best for medium to large production workloads (hundreds of GBs to multi-TBs/month), where query performance, structured metadata, and long-term scalability matter.
INGESTER
EFS
Logs to WAL for durability
S3
Logs (chunks) are buffered in memory.
TSDB shipper uploads blocks from EFS
Loki fetches blocks from S3. It can cache them on EFS, so repeated queries are faster
Local staging files can be deleted once upload succeeds.
Chunks flushed → written as blocks to EFS
04
Deploying a Loki instance, an OpenTelemetry collector instance, and an instance of Grafana to AWS using ECS and Fargate
04
For Loki we need the following Dockerfile:
�# Use the official Loki image as the base
FROM grafana/loki:3.4.5
# Copy the Loki configuration file into the container
COPY loki-config.yaml /etc/loki/loki-config.yaml
# Define Loki's startup command
CMD ["-config.file=/etc/loki/loki-config.yaml"]
Note:
04
And a file called loki-config.yaml that contains the Loki configuration and is the following:
BoltDB-shipper with EFS (v11 storage schema)
BoltDB-shipper with EFS and S3 (v11 storage schema)
TSDB with EFS and S3 (v13 storage schema)
04
These are the steps to upload the docker image to ECR:
Example commands for MacOS:
Install aws cli:
Login to AWS from the command line:
(registry uri example : 897729145798.dkr.ecr.eu-south-1.amazonaws.com)
We navigate to the folder that the Dockerfile exists and run:
Once its uploaded, we create a task definition for the Loki service:
04
04
Now that the task definition is in place, we’ll create a Loki service, by using the task definition and by enabling the service discovery.
For example, if we use “loki-serv-discovery-test” as a service discovery name and “namespace-test” as a namespace we can then use http://loki-serv-discovery-test.namespace-test:3100 to access Loki from another service in the VPC (OTel collector, Grafana)
This is done to have a DNS name that the services can use, instead of complicated and costly solutions like assigning a load balancer and a target group to the service and then a DNS name to the load balancer.
04
We’ll continue the process for the Otel collector. This is the Dockerfile we’ll use:
# Start with the official OpenTelemetry Collector Contrib base image
FROM otel/opentelemetry-collector-contrib:0.121.0
# Copy custom configuration file into the container
COPY ./otel-collector/otel-collector-config.yml /etc/otel/config.yaml
# Expose the necessary ports
EXPOSE 8889 13133 55679 4317
# Set the default command to load the custom configuration
CMD ["--config", "/etc/otel/config.yaml"]
04
And this is an example configuration for the OTel collector named otel-collector-config.yml:
Native OTLP ingestion
(TSDB schema v13)
Loki native (non-OTLP) ingestion (BoltDB-shipper schema v11)
Once its uploaded, we create a task definition for the OTel service (screenshot on the right):
04
Now, if we are still logged in to AWS, we can try to push the image directly to ECR:
Similarly to Loki, now that the task definition is in place, we’ll create an OTel collector service, by using the task definition and by enabling the service discovery.
04
Now we’ll push the Grafana image to AWS ECR. This is the Dockerfile:
FROM grafana/grafana:12.1.0
# Copy the custom datasource configuration file into the container
COPY grafana-datasources.yml /etc/grafana/provisioning/datasources/datasources.yml
# Expose Grafana's default port
EXPOSE 3000
# Optionally set environment variables to enable anonymous access:
# You can also define this in your ECS task definition.
ENV GF_AUTH_ANONYMOUS_ENABLED=”false"
ENV GF_AUTH_ANONYMOUS_ORG_ROLE="Admin"
# Use the default command for Grafana
CMD ["grafana-server"]
04
And this is the grafana-datasources.yml file needed by Grafana:�
Again, if we are still logged in to AWS, we can try to push the image directly to ECR:
04
The task definition for the Grafana service:
Similarly to Loki and the OTel Collector, we’ll create the Grafana service, by using the task definition and by enabling the service discovery.
04
We’ll now go to our security group and add the entry for port 3000. This is what you should see now in the security group’s inbound rules:
Please bear in mind that the settings used throughout this course are for development purposes. In a production env, Grafana should run in a subdomain with an SSL certificate in place, Loki should run in different security group allowing traffic from OTel collector and Grafana, etc.
05
Setting up a .Net API to write logs to our logging system
05
Simple example of setting up our .Net API to write logs to our logging system through the Serilog OpenTelemetry sink.
This is how our .csproj looks like:
�
05
And this is how Program.cs looks so that Serilog is setup as the logging provider, while configuring the OpenTelemetry sink with OtlpProtocol.Grpc so that our logs are being sent to the Otel collector through GRPC:
05
This is a screenshot from Grafana displaying test logs (navigate to the Fargate task’s public IP at port 3000):
06
Short mention of the steps required to include tracing and metrics for a full observability suite
06
To build a complete observability suite logs are not enough. We need traces and metrics.
Why do we need traces and how do they help?
We want to correlate logs under the same transaction.
Let’s think of what is Instrumentation: adding code or libraries to your application so it actually produces telemetry data (logs, traces, metrics) in a standard format (e.g., OTLP) that your observability stack can consume.
If you add the right packages and configure OTel in your .NET app, many spans are recorded automatically without changing your business code.
06
Why do we need metrics?�Metrics are numeric measurements that can be aggregated over time (e.g., counters, gauges, histograms).�They let us answer questions like:
Using OpenTelemetry for metrics ensures each measurement is enriched with resource attributes (service name, environment, version), making them easier to correlate with logs and traces.
�With the code below in Program.cs and two instrumentation packages (OpenTelemetry.Instrumentation.AspNetCore, OpenTelemetry.Instrumentation.Http), we can start exporting traces and HTTP-related metrics to the OpenTelemetry Collector:
From then on, we can use Jaeger or Tempo for traces and Prometheus for metrics. We will need though to make additions to the OTel collector configuration and task definition and also to Grafana, in order to have access to this information. The Jaeger traces are also visible from Jaeger’s own UI.
THANK YOU