1 of 49

Mastering Incident Reporting in the WebPKI

Driving Continuous Improvement Through Effective Incident Reporting

Ryan Hurst, CEO at Peculiar Ventures

August , 2024

2 of 49

I am most known for my work related to encryption on the web but over the last 30 years I have worked in a number of other areas like networking, operating system security, authentication, applied research, and standards.

Throughout my career, one of my primary goals has been to improve security for the next generation.

Some former employers include:

3 of 49

The beginning is the most important part of the work.

Plato

4 of 49

Why Browsers Care…

  • When a browser trusts a CA, it delegates some of the trust its users have placed in it to that CA.
  • When CAs that the browser trusts fail, it can harm the browser's credibility.
  • When browsers act responsibly concerning the trust placed in them, they help avoid legal consequences.
  • If the web is perceived as unreliable and insecure, browsers will see an attrition of users and, in turn, reduced adoption.
  • There are 86* CAs and Incident Reports are one of the few ways a browser gets to see what happens on the inside of these CAs.

* This is the current count of organizations in the Microsoft Root Program

5 of 49

Running a Root Program Is Hard

  • Root programs aspire to be objective and fair to support the global and open nature of the web.
  • This is why we have 86 WebPKI CAs in aggregate across the web.
  • These teams are responsible for policing the CAs often are as small as 1-3 people.
  • The reality is that only 7 CAs cover 99% of all certificates in use.
  • The remaining 79 CAs represent a huge attack surface that there is little visibility into and incident reports help with that.

6 of 49

In incident reporting, remember that root program managers often work in constant crisis management, so our reports must make it easy for them to get the necessary details.

7 of 49

Trust takes years to build, seconds to break, and forever to repair.

Unknown

8 of 49

Why Do We Do Public Incident Reporting?

  • Quick, comprehensive, clear and transparent incident reporting helps keep the web secure.
  • It establishes an incentive system to do the practice work that enables an organization to respond competently.
  • By requiring timely disclosures, subscribers and relying parties are better informed about the risks and impacts of a given certificate authority.
  • Enables the rest of the web to learn from the mistakes and practices of other CAs strengthening the web overall.
  • Help set benchmarks for incident response and prevention across the industry.
  • Makes it possible for all root programs to have access to the same information when evaluating their response to an incident.
  • Helps maintain trust in the web overall by making the security assumptions we make daily on it functioning correctly transparent.
  • It holds the reporting entity accountable for the resolution and prevention of future incidents.
  • Enables those that are exposed to the consequences of the incident to have an opportunity to engage and understand.

9 of 49

If you're doing it right, you make the web safer and provide more value than the risk you represent.

10 of 49

The single biggest problem in communication is the illusion that it has taken place.

George Bernard Shaw

11 of 49

Root Program Managers

Root program managers are the guardians of trust in the WebPKI. They scrutinize incident reports to assess CAs' compliance, security practices, and commitment to improvement. Their focus is on maintaining the integrity of the entire ecosystem, looking for patterns that might indicate systemic issues.

and more…

12 of 49

CA Customers

Certificate Authority customers range from small website owners to large enterprises. They read incident reports to evaluate the reliability and trustworthiness of their CA. Their primary concerns are the potential impact on their operations and the CA's ability to handle issues effectively.

and more…

13 of 49

Other CA Operators

Other CA operators closely monitor incident reports to stay informed about industry trends and avoid similar mistakes. They're looking for insights into common failures, and best practices that can improve their own processes and systems.

and more…

14 of 49

Relying Parties

DevOps engineers, security engineers, and PKI operators in organizations follow developments in the WebPKI as part of their job to understand potential impacts on their businesses, so they can adjust their practices accordingly. They are particularly interested in whether they need to switch to more responsible certificate providers.

15 of 49

Open Source Developers

Developers of open source projects related to PKI and web security analyze incident reports to ensure their software remains robust and secure. They look for insights that might necessitate changes in their code, and try to help hold CAs accountable for proper root cause analysis.

16 of 49

Tech Journalists

Tech journalists covering cybersecurity and internet infrastructure dive into incident reports to uncover newsworthy stories. They're interested in translating the issue into accessible narratives that drive clicks, focusing on the broader implications for internet users.

17 of 49

Understand Your Audience

  • Remember the reader will often take an adversarial approach and will find issues even if you do not.
  • Be sure to demonstrate empathy and humility.
  • Do not assume that anyone reading the report understands your organizational structure or how your systems work.
  • Do not presume that anyone reading the report understands the rules or your interpretation of them.
  • Be sure to look holistically at all related issues both open and closed, from you and other CAs before finalizing.
  • Do not simply state conclusions, always show how you came to those conclusions.

  • Make sure you start the timeline at the beginning of the problem, not just from the point of notification.
  • Make sure that you have clearly identified the true root cause and not just a symptom.
  • Make sure you talk to past commitments that might be related to the incidents because they will.
  • Read your wording carefully because incident reports are forever and what you say may end up in the press.
  • How you respond will impact your larger business not just the certificate issuance offerings.

18 of 49

The greatest enemy of knowledge is not ignorance, it is the illusion of knowledge.

Stephen Hawking

19 of 49

False. When done correctly, incident reports help improve the WebPKI, allowing us to enhance our own processes and highlight areas for improvement.

Common Misconceptions

Incident Reports are Bad

20 of 49

False. Incident reporting is a part of a continuous improvement process.

When done correctly, you are always preparing for the next incident and responding to the current one.

When combined with tabletop exercises, tooling, analytics, automation, and product roadmap additions incident reports help make it possible to both reduce the chance of a future incident and help ensure you are ready for the next one.

Common Misconceptions

Incident Report Reporting is Over When The Incident Is Closed.

21 of 49

False. Incident reporting is a family affair, compliance, engineering, operations, product, and leadership all have a hand to play both in preparing for incidents and managing through them. No one discipline is enough.

Common Misconceptions

Incident Reporting is a Compliance Function.

22 of 49

False. Most incidents are a function of engineering failures, organizational culture, or reliance on people rather than tooling to accomplish complicated and repetitive tasks.

Common Misconceptions

If we have an incident it is a compliance failure.

23 of 49

False. Incident responses are part of operating a production service, and the WebPKI is no exception. Only the most extreme single incidents result in distrust, such the CAs that have willfully violated trust of the web by issuing MiTM certificates.

Common Misconceptions

We Will be Distrusted if We Have an Incident.

24 of 49

False. Incident reports often reveal unforeseen weaknesses or systemic issues that are not the result of negligence but rather the complexity of technology and interdependencies within systems this often represents an opportunity to simplify.

Common Misconceptions

Incident Reports Always Point to Negligence.

25 of 49

False. Transparently handling and reporting incidents can actually enhance an organization's reputation by demonstrating commitment to continuous improvement and honesty.

Common Misconceptions

Incident Reports Will Always Damage Reputation.

26 of 49

False. While timely response is important, effective incident management also involves thorough analysis to prevent future occurrences, maintaining communication with stakeholders, and improving processes continuously.

Common Misconceptions

Effective Incident Management is Solely About Quick Resolution.

27 of 49

The only real mistake is the one from which we learn nothing.

Henry Ford

28 of 49

Browser distrust events of WebPKI Certificate Authorities occur on average approximately every 1.23 years.

29 of 49

Reasons Cited In Distrust Events

If we look at these events we see some common themes.

30 of 49

New incidents are filed every week.

Each one of these incidents is an opportunity to reassess our practices and systems, ensuring we do not repeat the same mistakes, while also advancing the entire WebPKI ecosystem's resilience and security.

31 of 49

The more you sweat in practice, the less you bleed in battle.

Richard Marcinko

32 of 49

  • Discovered a bug in internally through through compensating controls.
  • Acted swiftly by revoking over 1.7 million certificates within hours, prioritizing certificates still in use.
  • Demonstrated effective crisis management by quickly identifying the impact and guiding the revocation process.
  • Filed a delayed revocation incident explaining the delay in addressing remaining certificates, maintaining transparency.

Let’s Encrypt

DigiCert

  • Issue identified externally, indicating gaps in internal compliance and engineering processes.
  • Faced challenges in managing mass revocations, reflecting a lack of preparedness.
  • Struggled with scope determination, leading to a 5-day delay, indicating a lack of preparedness.
  • Failed to ensure customers were prepared and failed to adequate communicate to customers, , indicating a lack of preparedness.

33 of 49

Let’s Encrypt

DigiCert

Burndown of impacted certificates

Temporary restraining order resulting from not having done the work to minimize scope and impact.

34 of 49

  1. Preparation is Key: Automated systems, and tooling to support and predefined response strategies are crucial. For effective incident management, CAs should develop and test incident response plans regularly to ensure rapid action during a crisis.
  2. Transparency Builds Trust: During crises, maintaining open and regular communication is essential. CAs should ensure that their communication strategies are clear and consistent to build and maintain trust with stakeholders and the community.
  3. Learn from Others: Assign teams and individuals to conduct regular reviews of both historical and current incidents. Have them present these findings to the organization and rotate this responsibility across different disciplines to ensure knowledge is shared.
  4. Rely on Tools and Data, Not Just People: Use automated tools and data-driven strategies to ensure standardized and reliable incident responses, reducing dependence on manual judgment and minimizing errors.

Key Lessons

35 of 49

We are what we repeatedly do. Excellence, then, is not an act, but a habit.

Aristotle

36 of 49

37 of 49

Everything you do and don’t do sends a message.

Mike Hurst

38 of 49

  • CAs signed on to meet the browsers' requirements, and it is their responsibility to stay current with these requirements as they evolve.
  • Changes to CA/Browser Forum Requirements only occur if the CAs also vote for them. If a CA did not think the changes should have passed, they should have proposed alternatives and made a better case at that time.
  • If a CA was unable to convince enough CAs to join them in dissent, then when it passes, a responsible CA would make the necessary changes to meet the requirements.
  • If it was an oversight in the review of the proposal, the question becomes why the CA did not have enough, or the right, people involved to catch the issue.
  • Despite all of the above, the time to re-address the issue is once the incident is closed, not during the incident.

Common Mistakes

#1 : Arguing that the rules should change during an incident.

39 of 49

Common Mistakes

#2 : Claiming the issue is non-security relevant as an excuse.

  • CAs seldom have enough information to accurately assess the impact on subscribers or relying parties, and those involved in incident responses at CAs often mis-assess the impact and severity. For example:
    • The CA may not know that an attacker has necessary resources or knowledge to take advantage of of the weakness at the time of its finding.
  • The requirements are always requirements; there is no “you must do this unless it isn’t a security issue” exception in WebTrust, IETF, or root program policy.

40 of 49

  • Root programs treat all parties similarly; if they grant one CA permission to do something, it is expected that they should grant every other CA the same permission under the same circumstances.
  • Root programs do not create all the rules, nor do they have the authority to override the rules set by others, even if they wish to.
  • Lowering the standards for one CA puts the trust in the WebPKI and the browser in question.

Common Mistakes

#3 : Asking root programs for permission fail to meet the requirements.

41 of 49

  • The decision to distrust is a combination of many different root programs' views; if you are distrusted by one and not by others, your CA's ability to effectively service your customers is almost always harmed.
  • The decision to distrust a CA for incident handling will ultimately go through many individuals in each root program, and through this process, an initial disposition from one individual may change after a more complete analysis.
  • Having such conversations cheats the community of the opportunity to learn from the incident and may result in the issue being raised again later, potentially leading to a different disposition at that time.
  • Root programs are not your compliance team, nor are they your auditor, or the community as a whole. Bring your questions to the right audience.

Common Mistakes

#4 : Asking a root program in private if the CAs understanding is correct.

42 of 49

  • To make it easy to ensure incidents are analyzed consistently a common reporting format is used.
  • This common format makes it easier for the community and the root programs to quickly analyze a incident and understand what has happened.
  • The common format also makes it easy to report compare once incident to another incident report.
  • Using formats like PDF, Word Documents and others make it hard for the community to uniformly analyze the whole set of statements made by the CA.

Common Mistakes

#5 : Not following the standard reporting templates

43 of 49

  • Many bad incident reports will identify the root cause as things like “we missed this in our review”, “compliance failed to detect the issue” or “we had a bug” without asking the question of why did that happen and what can we do to address it.
  • The associated investigations often fail to explore all potential contributing factors, such as weak engineering practices, lack of automation, reliance on manual processes, procedural weaknesses, staffing limitations and organizational culture.

Common Mistakes

#6 : Failing to identify the true root cause of an issue

44 of 49

  • Failing to promptly notify affected parties can exacerbate the impact of security incidents.
  • Not providing all relevant information about the incident on purposefully or through lack of preparation.

Common Mistakes

#7 : Insufficient Communication with Affected Parties

45 of 49

  • Failing to do the necessary preparatory operational and engineering work to ensure that prescribed timelines are met effectively.
  • For example:
    • Not understanding how to proactively reach customers.
    • Not proactively measuring if certificates are in use so that revocations can be prioritized based on impact.
    • Lacking tooling to automate incident response-related tasks.
    • Failing to make engineering investments to ensure timely revocation and re-issuance.

Common Mistakes

#8 : Non-Compliance with Prescribed Timelines

46 of 49

  • Repeating the same or similar mistakes of other CAs or worse, their own, indicates a failure to implement effective lessons learned from past incidents.
  • Not showing continuous improvement in handling and preventing incidents demonstrates a lack of commitment to maintaining trust and security.

Common Mistakes

#9 : Failure to Learn from Past Incidents

47 of 49

  • When a CA commits to completing engineering or operational changes to address a class of issues and fails to do so, it not only reflects a disrespect for the root programs and the trust placed in the CAs, it also signals weak compliance and engineering practices.

Common Mistakes

#10 : Failure to do the work committed in past incidents

48 of 49

If your Incident Response process is not “boring” yet, you have not mastered it.

Proper preparation prevents poor performance - Charlie Batch.

Want to learn more about this topic?

49 of 49

Other Thoughts

  • Your internal review processes should capture the necessary evidence to quickly and accurately prepare your public incident response. For example, when reviewing a CPR and assessing its significance, that analysis should be documented in case it's needed later.
  • You should regularly rely on your auditors to verify your internal assessments of correctness, rather than relying on individual personal judgments.
  • Regular tabletop exercises for theoretical incidents are crucial. Utilize your playbooks and tools to check if you can meet your SLA obligations based on your training, tools, and processes. This not only trains your team but also drives continuous improvement.
  • Beyond tabletop exercises, conduct regular training sessions and simulations for all employees to familiarize them with incident response procedures and encourage a culture of awareness of the requirements and processes.
  • A post-mortem process for all incident reports is essential to understand what went well and what didn’t. Use these insights to revise training, policies, procedures, tools, and products.
  • You should be doing cross discipline reviews of CA failures, public incident responses, and even bringing in individuals from other CAs to talk about how they do things so you can learn from them.