1 of 22

A unified topology system for �a large scale, heterogeneous and dynamic computing infrastructure

Alexey Anisenkov (BINP)

on behalf of CRIC team�

1

CHEP 2019, Australia, Adelaida, 7 Nov 2019

2 of 22

CRIC: a high-level information middleware to configure computing environment and describes resources as VOs need

2

Alexey Anisenkov, CHEP-2019

Rucio

Pilots

OIM

Open�LDAP BDII

REBUS

GOCDB

Other sources

- Low-level infosys

- Data providers

- Service discovery

WLCG Central �ops

GlideinWMS

Phedex

WLCG squids

And many others…

Resources description

Experiment applications

3 of 22

CRIC mission: link Resources & VOs together

  • Consolidate (WLCG)** topology information of a large scale computing infrastructure
  • Facilitate distributed computing operations for (LHC)** Experiments

Key functional capabilities of CRIC information concept:

      • Built-in information Model(s) for Resource descriptions
      • Clear distinction between (physical) resources provided by (Sites) and how they are used by Experiment(s)
      • Built-in aggregation and validation of data collected from various low-level information providers (sources)
      • Ability to extend and complement Information Model(s) with Experiment(s) specific data structures. Plugin based approach allows VOs to inherit shared functionality and address custom requirements
      • Flexibility to address technology evolution and changes in the VO Computing models and applications
      • Experiment-oriented but still Experiment-independent IS framework

3

Alexey Anisenkov, CHEP-2019

** CRIC is not coupled to WLCG, it could be applied beyond LHC experiments

4 of 22

CRIC Architecture: examples of shared features out of the box

  • Plugin based: VO can configure default behaviour
  • Base implementation for the Resource/Topology description
  • REST API data export (filters, presets, various output formats)
  • Shared engine/widgets for WebUI (downtime calendars, table view, tree view, inline editors, etc..)
  • Enhanced Authorization (CERN SSO, SSL, paswd based;local accounts)
  • Enhanced Authentication (instance specific permissions, groups, roles, etc, map permissions to e-groups, fetch info from ext sources)
  • Detailed History of Changes (who, when, how interacted with object)

4

CORE

CORE

CORE

CMS CRIC

ATLAS CRIC

WLCG CRIC�(serves ALICE, LHCb, overall view)

REST API

WebUI

SiteDB

REBUS��

Other Sources

BDII

GlideinWMS configs

COMPASS CRIC

CORE

CORE

Generic CRIC

VOs can use shared WLCG CRIC instance as the master source or optionally fetch data directly from low-level info providers

5 of 22

Few Implementation details: Web2.0 based

5

Alexey Anisenkov, CHEP-2019

  • Apache/WSGI + Python + Django framework as server backend
  • Independent database backends (Oracle, MySQL, postgres, etc)
  • Web Services technologies �(REST API, WebUI, widgets)
  • Bootstrap framework as HTML/CSS/JS client frontend�(responsive, interactive, mobile-friendly)
  • Client AJAX, JQuery plugins, own shared widgets (datatables, treeview, calendar views, inline editors ..)
  • Plugin based approach �(shareable applications in “core” re-used by many components)

6 of 22

CMS CRIC

In production for CMS Community since the end of 2018

  • As a replacement of CMS SiteDB (currently in read-only mode, to be retired soon)
  • Basic topology description (sites, processing unit names translation, etc)
  • User Groups, roles and privileges used by CMS services (CRAB3, Phedex, WMAgent, cmsweb, etc)
  • CMS User description required by CMS applications
  • Job submission configuration (GlideinWMS) �(see the talk “Exploiting CRIC to streamline the configuration management of glideinWMS factories for CMS support” by James Letts)
  • Various REST API export, including integrated CORE functionality and CMS specific objects of Information Model (CMSSite, ComputeUnit, ComputeResources, GlideInEntry, ..) �(ongoing, step by step integration into production)

6

Alexey Anisenkov, CHEP-2019

7 of 22

WLCG CRIC

Dedicated CRIC instance for central WLCG operations

  • Single entry point for complete WLCG topology description and service configurations for the all 4 LHC experiments
  • Main info provider for cross-experiment tools: �WLCG Accounting, Monitoring, Service Availability, Test submission systems,..
  • Federation Pledges management and topology export (REBUS replacement)
  • VOFeed XML generation (ALICE, LHCb)
  • Management of VO Pledge Requirements
  • Tracking of various Task Forces and Migration activities
  • WLCG Accounting data validation (storage space and CPU capacity from WSSA)
  • WLCG Accounting Report generation interfaces
  • In production since Summer 2019. Site Admins started to validate and update information

7

Alexey Anisenkov, CHEP-2019

8 of 22

ATLAS CRIC (ongoing)

Moving from AGIS to ATLAS CRIC has started

In fact CRIC is the evolution of AGIS (ATLAS Grid Information system) - completely refactored code, inherits all AGIS features

8

Alexey Anisenkov, CHEP-2019

    • All API export should be implemented in backward compatible way (no significant changes expected for exist ATLAS clients)
    • CRIC will sync data from AGIS.and advertise own API for the integration period (AGIS is still a master for given data block once “edit” functionality is migrated into CRIC)
    • Redirect back to AGIS for not covered functionality
    • Once a functional block is migrated to CRIC, disable modification in AGIS and redirect to CRIC for the migrated part

Tentative Plan for the migration:

  • Basic topology, Blacklisting API, PanDA SW releases -- Oct/Nov 2019
  • Moving slightly dependent data and settings -- Nov 2019
  • Complete PQ+RSE management within CRIC -- Beginning of 2020

9 of 22

DOMA CRIC

Dedicated CRIC instance for Rucio TPC tests� and DOMA related activities

  • Provides Storage description and related data structures for �Experiment agnostic DOMA Rucio instance
  • In close cooperation with Rucio experts to polish RSE related models and CRIC interfaces in order to provide appropriate API export for Rucio clients (probes)
  • Rucio team has tested RSE configuration coming from DOMA CRIC with Rucio ESCAPE instance. All works well. Look forward for the next integration steps.

  • Once CRIC and Rucio integration will be completely tested and evaluated, developed CRIC models and interfaces within DOMA CRIC will be shared with other plugins (ATLAS CRIC, CMS CRIC)

9

Alexey Anisenkov, CHEP-2019

10 of 22

COMPASS CRIC

Dedicated CRIC instance for COMPASS Experiment � at CERN SPS

  • COMPASS Distributed Computing Environment is very similar to one used by ATLAS (Computing Model, PanDA WMS, PanDA Pilot, ..)

  • First COMPASS CRIC prototype has been evaluated, tested and initially integrated into COMPASS environment.
  • Today COMPASS upgrades components of WMS (PanDA Pilot, Harvester migration) that in particular requires updates from COMPASS CRIC side

  • Current status: integration step, to be released soon.

10

Alexey Anisenkov, CHEP-2019

11 of 22

Lightweight CRIC plugin (ongoing)

Universal topology description of generic distributed infrastructure

  • Enables all CRIC features but with simplified Computing Model description
  • Basic models for Compute and Storage Resources (StorageUnit+StorageResource, ComputeUnit+ComputeResource)
  • Completely CERN-independent
  • Standalone distribution (via images), not coupled to CERN Openstack deployment infrastructure
  • Suitable for small VOs or Experiments beyond LHC

Requested by Experiments at JINR (NICA and beyond)�as the Information component�for the Unified Resource Management System

11

12 of 22

Collaborative support model

  • CRIC project is a joint effort between CERN (IT-WLCG and the Experiments), Novosibirsk (BINP), DUBNA (JINR) and other institutions.
  • CRIC core services are supported by the CRIC team
  • For the CRIC plugins we expect active participation from the Experiments mainly for functional requirements gathering, light developments and polishing of affected experiment-specific features

  • Collaborators from the Experiments are welcome!
    • Developments, use cases gathering, sustainable strategy planning...
    • Please join and get involved to deliver together powerful and sustainable experiment-oriented CRIC tools!

12

Alexey Anisenkov, CHEP-2019

13 of 22

Conclusion. CRIC family

  • CRIC offers a common framework describing LHC Computing infrastructure with also an advanced functionality enabled to describe all necessary Experiment-specific configurations.
  • Ready to use now. Let’s play and/or contribute.

13

Alexey Anisenkov, CHEP-2019

in production

core-0.2.7

cms-0.2.6

CMS CRIC

core-0.3.0

wlcg-0.1.5

WLCG CRIC

core-0.3.0

doma-rc

DOMA CRIC

core-0.2.9

atlas-testbed

ATLAS CRIC

core-0.2.3

compass-rc

COMPASS CRIC

ongoing

lightweight CRIC

development

integration

14 of 22

Thank you for your attention!

Backup slides

14

Alexey Anisenkov, CHEP-2019

15 of 22

Distributed Computing Environment (Resources)

LHC Experiments rely on heterogeneous distributed computing

    • variety of computing resources involved �

    • variety of infrastructures and middleware providers ��

15

Alexey Anisenkov, CHEP-2019

Research granted access

Opportunistic backfilling

HPC

HPC

Grid

WLCG Pledged resources

Cloud

Volunteers

Rented, �on demand

Opportunistic

Opportunistic�backfilling

others ..

16 of 22

Distributed Computing Environment (Experiments)

  • Each Community uses and describes Resources in its own way

    • Computing Models are similar but still have different implementation
    • Various high level VO-specific frameworks & middleware services (e.g. for Data and Workflow management)
    • Cross experiments applications (monitoring, accounting, testing frameworks, resource usage descriptors, etc)
  • Apart from resources description, high level VO-oriented middleware services and applications also require the diversity of common configurations to be centrally stored and shared

16

Alexey Anisenkov, CHEP-2019

17 of 22

Authorization and Authentication (A&A)

  • CRIC supports enhanced Access controls and user Group management
  • Several Authentication methods enabled (SSO, SSL, Proxy cert, passwd)
  • Flexible utilisation of Permissions, Roles and Groups at various levels
  • Fine grained Auth checks at various levels (object, model, restricted instances, global permissions)
  • Ability to bootstrap DB (User info) from whatever external source� (CERN DB, Experiment DBs, config files, e-groups, VOMS roles, etc)

Each Experiment could configure own Data access policies!

17

User Profiles

SSO

SSL

VOMS

Local

User

Custom Groups

permissions

Roles

Group memberships

Auth Groups

Alexey Anisenkov, CHEP-2019

18 of 22

Example of A&A use-cases for different VOs

  • CMS uses CRIC not only to define access rights within the system, but also to control user privileges for CMS applications (CRAB, WMAgent, Phedex, etc…). Relies on CERN SSO and local authentication.
  • CMS enables instance specific permission checks (who is allowed to manage specific object(s), e.g. update CERN-PROD site, and/or all affected resources)
  • ATLAS considers simplified Auth concept based on user’s DNs

18

Experiment decides what elements should be used out of the CRIC box to implement own policies and follow own workflow.

Alexey Anisenkov, CHEP-2019

SiteDB�Groups

Groups membership

Auth Groups

per Facility Groups

Group Responsibilities

Auth Groups

CMS User

ATLAS User

Roles

Roles

Automatic sync with CERN User DB

19 of 22

Data export availability

  • Thanks to LB deployment (and appropriate DB backend) CRIC is able to sustain with high API load coming directly from worker nodes
  • For highest reliability, VO can enable dedicated CacheCheckers application to automatically cache and dump required API JSONs to special location (e.g. CERN EOS box)
  • Experts can subscribe to email notifications in case of JSON changes

19

Latest (stable) caches can be further mirrored to CVMFS (e.g. used by ATLAS) as (failover, cache) access to topology data from worker nodes

Alexey Anisenkov, CHEP-2019

20 of 22

Recent core developments (tableviews)

Inline editor for bulk updates

Flexible table view visualization

  • Sometimes it’s helpful to provide required set of parameters as a dedicated view (with ability to dynamically expose on fly all affected/predefined fields)
  • Activate inline modification for specific fields

20

Alexey Anisenkov, CHEP-2019

Update required fields for several objects,

Bulk save

Customize visualization on fly, prepare custom data view for share and distribution

21 of 22

Recent core developments (request user privileges)

ADMIN Privileges Request/approval wizard

  • VO can customize supported list of ADMIN Groups family �(Global ADMIN, Site Admin, Federation Admin, etc ..)

21

Alexey Anisenkov, CHEP-2019

Clicky-Clicky wizard to request privileges by user

ADMIN decides/grants only required perms.�Email notification sent on request/approval steps

22 of 22

Ongoing core developments

Moving to use-case oriented approach of updating data

  • Classical approach assumes to update some set of parameters for specific model of Information schema (="configure these variables for these resources")
  • In reality, typical use-case oriented (workflow) update
    • involves modification of some attributes of several affected models depending on user input
    • requires extra validation and conditional modification
    • user usually does not know which parameter is affected �(for example in ATLAS, PandaQueue model has ~ 100 parameters)
  • CRIC will provide a wizard-like workflow forms to process specific use-case for data modification
  • Example: I want to enable remote-io mode for jobs:
      • at which site?
        • for which type of jobs? (ANALY, PROD)
          • for which input storage?
            • for which type of access ? (LAN/WAN), ..

22

Alexey Anisenkov, CHEP-2019