Article

Project Management

Risk Management

September 29, 2023

Risk management for an educational institute website

DOI: 10.22167/2675-6528-20230069
E&S 2023,4: e20230069

Fabio Luis Gonzaga; Aline Bigaton

According to the Project Management Institute (PMI)[1], the global association of project managers, risk is any uncertain event or condition that, if it occurs, can cause positive or negative impacts on the project objective. Negative effects can represent threats; positive ones, opportunities. All projects have risks, and it is up to the project team members to identify them to minimize or avoid threats and maximize opportunities. Risk management is a process that helps identify, analyze, and plan responses, as well as monitor these risks[2].

A survey conducted by the Regional Center for Studies for the Development of the Information Society (Cetic.br) showed that 54% of Brazilian companies have a website, and 57% use the internet as a means to make sales[3]. Companies whose website is a critical component for the business should manage the risks related to the cost per downtime of the website in the same way they do to protect themselves from losses such as errors, omissions, labor claims, among others[4].

The web serves many organizations as an integral part of their daily operations to attract customers, communicate with suppliers, and generate revenue; as a result, the cost of failures in the availability of its services can be significant[5]. According to Plesky[6], downtime can be defined as the period in which the website is completely inaccessible or unable to perform its main functions, causing, even for a brief period, customer dissatisfaction. This can contribute to a drop in online search engine rankings and result in loss of customers and revenue, damaging brand reputation.

Between February 19 and 23, 2022, the websites of the Americanas S.A. group, responsible for the “marketplaces” of Lojas Americanas and Submarino, experienced access instabilities and were offline due to a security incident. According to the consultancy Economatica, it is estimated that the company suffered a loss of R$ 3.4 billion[7] as a result of the incident.

Websites can suffer from permanent, intermittent, or transient failures, affecting the platform as a whole or partially. Permanent failures persist until the problem is repaired; transient failures eventually disappear without apparent intervention; and intermittent failures are transient and occur occasionally, such as when there is a system overload. The causes of failures can be categorized as software, hardware, environmental, or human errors, in addition to security violations[5].

According to Pertet and Narasimhan[5], failure manifests itself through effects that are perceptible to users; examples include system exceptions and access violations, which produce error messages. Another type of failure indication is related to incorrect results, such as the display of a wrong or blank page, or the presence of incomplete information. Generally, this type of failure is only discovered after customer complaints. Failure can also manifest as slowness in performance, which can be attributed to various reasons, including server overload, process freezing, exhaustion of computational resources, or network congestion.

Given the relevance of the topic, a case study was conducted with the objective of developing a risk management project for the website of an educational institute. The research evaluated how risk management was carried out and the perception of the development team members regarding the importance of this management. After the diagnosis, the main risks to which websites are exposed, and particularly the website under study, were listed. Finally, a risk response plan was developed.

The study[8] was developed at a Brazilian institute of education and research, headquartered in the interior of São Paulo and with more than 500 collaborators. Among the services offered were undergraduate and postgraduate courses, technical analyses — mainly for the agribusiness sector —, management training, among others. The website chosen for the study was part of the institute’s product portfolio and advertised postgraduate courses in distance learning (EaD) on its page.

This website was selected as the study’s target due to its high visitor demand. More than 70% of all traffic originated from paid ads, making a highly available website critical. Upon being directed to the page, the visitor found information about the institution and each course; however, to enroll in one of them, the user was directed to another system, which collected personal and payment data.

An internal development team was responsible for developing new features, updating, and maintaining the website. This team included product professionals, user experience specialists, designers, developers, and data analysts. The identification of risks was limited to the website under study, excluding the system responsible for registration and payment, as well as marketing actions for lead generation (potential customers). The identified risks referred to failures that, if they occurred, could lead to downtime.

The methodology used in the present research was the case study, which has as a strategy to examine contemporary events with techniques used in historical research, but adding direct observation and a systematic series of interviews[9] as evidence sources. In the first stage, the technical characteristics of the site under study were mapped, with the purpose of understanding the relationship between the prevailing panorama at the time and the risk perception level of the evaluators regarding the most common threats in site projects. To aid in the diagnosis, the list of common risks from the “Project Management Body of Knowledge” (PMBOK)[1] was used, which contains items, actions, and points to be considered, based on historical information and accumulated knowledge from similar projects. Based on the result, interviews were conducted with team members and specialists for the qualification of risks, evaluating the probabilities and the impact in case of occurrence.

Salles Jr. et al.[8] describe that the probabilistic analysis of events is generally necessary when information about the process is lacking, which can lead to inaccuracies. Therefore, the more information obtained, the less uncertain the event will be. Thus, after identifying the common risks of the object of this study, an interview was conducted, through the application of a questionnaire, with members of the project team. The result was a risk response plan, proposing alternatives to prevent or reduce the business’s exposure to threats.

Throughout a project’s lifecycle, risks continue to emerge, and management processes must occur interactively; the project team needs to understand if the level of risk exposure is acceptable, and for this, the project’s size, complexity, and strategic importance must be considered[2].

The website object of the study used a cloud platform specialized in the technology and programming language for which it was developed. This platform used environment variables securely, preventing sensitive information from being exposed in the source code or transmitted through other means. Environment variables are important structures with dynamic values that are temporarily stored on the server and used during program execution; they support a continuous integration and delivery process, which automates and facilitates the deployment of new updates and ensures that the latest functional version remains available in case any error is detected, reducing the chances of a probable site failure[10].

Among the main features offered by the platform were: support for secure navigation through “hypertext transfer protocol secure (HTTPS)” — which functions as protection for communication and data transfer between the web browser and a website[11] — and the “secure socket layer (SSL)” security certificate, a protocol that requires the web server to have a digital certificate, functioning as a public key for data transfer through this connection[12]. The platform also offered features for mitigating distributed denial-of-service attacks, commonly called “distributed denial-of-service (DDoS)” attacks. DDoS attacks use thousands of infected machines to access a website simultaneously, causing system overload and, consequently, denial of service by the server[13].

The source code update was entirely done through code versioning processes, which allow for a reliable history of modifications and enable the restoration of a previous version in case a problem occurs. Access to the source code repository was restricted to members of the project development team.

According to Sutherland[14], the definition of done in agile methodology means that all conditions and criteria established by the team have been met, and tests have been performed. The website team in question used the following criteria as their definition of done for all update requests: 1) the increment must pass functional tests, and the tester must not be the same person who created the increment; 2) the increment must be reviewed by at least one other peer and only released after approval; 3) non-functional criteria such as accessibility, performance, and security must be considered during creation/development; 4) what was done must be documented; 5) acceptance criteria must be met; 6) the increment must align with user experience (UX) and user interface (UI) guidelines.

Table 1 presents a survey of common incidents that can cause website downtime, based on the history of actual occurrences suffered on various pages[5],[15]. There are other generic and common risks that can cause unavailability or affect the integrity of a website; however, for this study, only the risks to which it was believed the studied website could be exposed were considered. Each risk contains its own identifier and respective description.

Table 1. Identified risks

IDRisk
1Resource exhaustion: memory leak, resource-intensive processes,
causing page requests to time out, insufficient disk space
2Environmental factors: power outage, natural disasters
3Traffic overload: occurs mainly when there is an abnormal flow of visitors to the site,
either due to an advertising campaign or a possible cyber attack
4Attempts at attacks or malware: DDoS attacks, for example, can lead to a high demand
of traffic on the site maliciously, so that it becomes unavailable
5Encoding error: in the technology area, a bug is a coding error in a
computer program[10]. Unfortunately, a bug cannot always be detected
quickly
Source: Pertet and Narasimhan[5]; Jackson[15]

During the qualitative analysis, probabilities and impacts are evaluated, which aids in risk prioritization. After identifying the risks, it is possible to qualify them by assigning scores for the probability and impact of each event[9]. Tables 2 and 3, below, show the probability classification and the impact level identified for the studied site.

Table 2. Classification of the probability of occurrence

Probabilistic scaleDegree of probability of occurrence
0,1Very rare chance of occurrence
0,3Low probability of occurrence
0,5Moderate chance of occurrence
0,7High probability of occurrence
0,9Very high chance of occurrence
Source: Adapted from Salles Jr. et al.[8]

Table 3. Risk event impact classification

Probabilistic scaleDegree of impact
0,1Very low
0,3Low
0,5Moderate
0,7High
0,9Very high
Source: Adapted from Salles Jr. et al.[8]

The impacts were divided into three categories:

  • cost: financial impact generated for the company if the risk event occurred (for example, the loss of leads if there was a service unavailability that harmed or prevented registration);
  • strategic: a fact that could impact the performance of a marketing campaign or action or harm the brand’s image;
  • quality: a fact that caused a negative user experience, such as slowness, functionality failure, navigation problems, or page not found.

After identifying the main common risks to which the website was exposed, an interview was conducted through a questionnaire with team members, with seven respondents, one product manager and six developers. Each interviewee assigned a random value to the identified risks, following the probabilistic scale of 0.1 to 0.9 for the classification of probability of occurrence and risk event impact level, according to Tables 2 and 3, respectively.

Table 4 presents the risks prioritized by the group according to the highest calculated risk degree (column “risk”), which was obtained from the multiplication of the probability of occurrence (column “probability”) by the consolidated risk (column “consolidated”), this being derived from the highest value assigned as impact among the three categories (cost, strategic, and quality).

Table 4. Risk probability matrix

IDRisk eventProbabilityImpactConsolidatedRisk
CostStrategicQuality
1Resource
depletion
0,300,580,720,700,720,21
2Environmental
factors
0,210,640,580,640,640,13
3Traffic overload0,500,610,670,670,670,33
4Attempts
of attacks
or malwares
0,470,670,700,720,720,34
5Encoding error0,550,500,500,550,550,31
   Overall risk1,34
Source: Adapted from Salles Jr. et al.[8]

The overall project risk was normalized as recommended by Salles Jr. et al.[8], according to Equation 1:

(1)

where, RGP: is the general project risk; P: are the probabilities; I: are the impacts; n: is the quantity of identified risks (in this case, equal to 5); EP: the probability scale; and EI: the impact scale, both with values of 0.9.

As a result of the normalization within this study, an overall risk of 33% was obtained. The perception of high risk was due to the project’s sensitivity, reinforcing the relevance of the study of the topic and the need for the development of a risk management project with adequate response plans for each eventuality.

The graph below, represented by Figure 1, presents the risk matrix for each identified risk event. It helps to understand the degree of perception about the probability of each event occurring and its impact, should it occur. The matrix also helps to prioritize risks in a more coherent way, avoiding the use of assumptions.

Figure 1. Risk matrix
Source: Adapted from Salles Jr. et al.[8]

The risk response plan aims to assist the project team in foreseeing strategies and actions that should be executed to respond to threats identified during the risk identification process[8].

Based on the risk analysis performed in the study, a risk management project was developed to assist in the creation of a response plan, with the objectives of: offering guidelines for the adoption of good coding development and review practices; describing possible security issues; guiding the planning of the most appropriate server for the website; and recommending risk mitigation measures.

Errors or bugs can be caused by the absence or inefficiency in code review processes, failures or lack of security policies for implementation, or by the absence of good practices during development. In response, it is necessary for the development team to implement code review processes for each new update or maintenance in the project.

In a review, multiple aspects should be observed, such as code functionality, clarity of its writing, presence of useful comments, among other factors. Furthermore, the code review process can and should be used as part of the development team’s improvement and learning process[16].

It is recommended that the team involved in the project continue using the version control system (VCS). The adoption of this practice allows for a history of previous versions to be obtained with each modification, functioning as a backup that can be restored in case of any problem with a new version, thus generating a history of changes in the project throughout its lifecycle and facilitating collaborative work among developers on the same project[17].

Websites are constantly susceptible to cyberattacks. Among the main types are: DDoS attack; “man-in-the-middle (MitM)”; “phishing”; “spear phishing”; “drive-by”; brute force; “SQL injection”; “malware”; and “cross-site scripting attack (XSS)”. Table 5 details each type of attack.

Table 5. Common types of cyberattacks on websites

NameDescription
Brute force attackIn this attack method, hackers use algorithms that perform thousands
of combinations to discover login and password
“Cursory”Method used for “malware” dissemination; in this type of attack,
hackers look for insecure websites and inject malicious “scripts” (codes)
in the HTTP protocol, for example
“Malware”Also known as malicious software, it is a program or file
intentionally harmful to the device
“MitM”Abbreviation from English for “man-in-the-middle”. In this type of attack,
the attacker inserts themselves between the communication of a client and a server.
Examples include session hijacking and IP spoofing
“Phishing” and “spear phishing”It is the attack modality used to send emails impersonating trusted sources, with the objective of obtaining privileged information or influencing the user to perform some action; another technique used by scammers is cloning legitimate websites with the intention of obtaining personal information or access credentials to the true website
“SQL injection”English term for “structured query language (SQL)” injection.
Language used in databases, it is a common type of attack in
which SQL commands are inserted, which can range from capturing, inserting,
altering or deleting customer information from the database
“Cross-site scripting (XSS)”It is a type of attack used mainly to exploit vulnerabilities
that allow, while the victim browses the website, the hacker to capture
information stored in the browser, screenshots, keystrokes
, network information, or even control the victim’s machine
Source: Melnick [13] ; Technological target [11]

Protecting oneself from these and other attacks requires understanding the concept of offense. For Melnick[13], measures aimed at mitigating threats can vary, but there are basic approaches to keeping systems and virus databases updated, good team training, correct firewall configuration, strong passwords, and regular backups. It is important for the website to have a security certificate, even when it does not handle confidential information. Navigation via the HTTPS protocol, in addition to protecting against misuse of the site, is a mandatory requirement for many technology resources[18].

Server downtimes can occur due to lack of maintenance or unexpected defects. It is important to choose a service of recognized quality, such as the content delivery network (CDN), which, according to Plesky[6], can be a resource adopted to create a layer between the server where the website is hosted and the user. The service enables content delivery through geographically distributed servers, using content caching systems, which avoids website unavailability when the server is unavailable for a short period. The CDN also helps prevent malicious “bots” (robots) from accessing the site, by filtering traffic and preventing server overload.

The hosting platform of the studied website offered the CDN service active by default, without the need for additional configuration or contracting. Plesky[6] highlights that domain expiration can also cause website unavailability. The author therefore recommends purchasing for long periods or automatic renewal, as is the case with the studied website’s domain.

Downtime can also be planned for normal IT infrastructure operations, such as backups, maintenance activities, or system patches. However, planned downtime is one of the main factors causing incidents, such as, for example, when a backup takes longer than planned[5].

The cloud platform of the website object of this research had “uptime” (time the server is operationally available[11]) of 99.99%, guaranteed by the provider company. The platform also had a page for checking service status and support channels for reporting failure occurrences.

It was possible to perceive, through this work, that the studied platform had a service adequate to the project’s scope; nevertheless, it would be advisable to evaluate a second option for rapid deployment on alternative hosting, in case of environmental factors that may make the site unavailable for long periods. Having a risk response plan does not guarantee that the site will be completely secure, but it helps the team to proactively think about continuous improvement processes.

As a recommendation to complement this study, periodic security audits on the website are suggested to find possible vulnerabilities and mitigate the possibility of cyberattacks. It was concluded that, for the studied website, the probability of occurrence of the listed events was low, mainly due to the mitigation actions already adopted by the development team; however, the impacts can be high if they do occur.

References

[1] Instituto de Gerenciamento de Projetos (PMI). 2021. O padrão para gerenciamento de projetos e um guia para o corpo de conhecimento em gerenciamento de projetos (guia PMBOK). 7ed. Project Management Institute, Newtown Square, PA, EUA.

[2]Project Management Institute (PMI). 2017. Um guia do conhecimento em gerenciamento de projetos. 6ed. Project Management Institute, Newtown Square, PA, EUA.

[3] Empresa Brasil de Comunicação (EBC). 2020. Mais da metade das empresas brasileiras usam internet para vender e 78% estão nas redes sociais. Disponível em: <https://agenciabrasil.ebc.com.br/radioagencia-nacional/acervo/economia/audio/2020-04/mais-da-metade-das-empresas-brasileiras-usam-internet-para-vender-e-78-estao/>. Acesso em: 10 abr. 2022.

[4] Pagely. 2015. Risk Mitigation and the True Cost of Website Downtime. Disponível em: <https://pagely.com/blog/risk-mitigation-and-the-true-cost-of-website-downtime/>. Acesso em: 10 abr. 2022.

[5] Pertet S.; Narasimhan P. 2005. Causes of Failure in Web Applications. Technical Report, Parallel Data Laboratory, Carnegie Mellon University, Pittsburg, PA, EUA.

[6] Plesky E. 2021. What is Website Downtime and Why Should You Take it Seriously? Disponível em: <https://www.plesk.com/blog/various/what-is-website-downtime-and-why-should-you-take-it-seriously/>. Acesso em: 10 abr. 2022.

[7] G1. Tecnologia. Sites de Americanas e Submarino voltam a funcionar após três dias fora do ar. 2022. Disponível em: <https://g1.globo.com/tecnologia/noticia/2022/02/23/americanas-tem-site-reestabelecido-depois-de-quatro-dias-fora-do-ar.ghtml>. Acesso em: 10 abr. 2022.

[8] Salles Jr. C.A.C.; Soler A.M.; Valle J.A.S.; Rabechini Jr. R. 2006. Gerenciamento de riscos em projetos. FGV Editora, Rio de Janeiro, RJ, Brasil.

[9] Yin R.K. 2001. Estudo de caso: Planejamento e métodos. 2ed. Bookman, Porto Alegre, RS, Brasil.

[10] Oliveira M.E.; Zuccherelli M.F.L.; Libera G.P.D.; Oliveira R.L.Z.; Tech A.R.B. 2020. Introdução à robótica educacional com Arduíno – hands on!: iniciante. Faculdade de Zootecnia e Engenharia de Alimentos (FZEA/USP), Pirassununga, SP, Brasil. DOI: 10.11606/9786587023052.

[11] Tech Target. 2022. Computer Glossary, Computer Terms – Technology Definitions and Cheat Sheets from WhatIs.com – The Tech Dictionary and IT Encyclopedia. Disponível em: <https://www.techtarget.com/whatis>. Acesso em: 13 set. 2022.

[12] Bhiogade, M.S. 2001. Secure Socket Layer. In: Anais do Computer Science and Information Technology Education Conference; 2002; Cork, Irlanda. p. 85-90. DOI: 10.2139/ssrn.291499.

[13] Melnick J. 2018. Top 10 Most Common Types of Cyber Attacks. Disponível em: <https://blog.netwrix.com/2018/05/15/top-10-most-common-types-of-cyber-attacks/>. Acesso em: 25 jul. 2022.

[14] Sutherland J. 2014. Scrum: a arte de fazer o dobro do trabalho na metade do tempo. LeYa, São Paulo, SP, Brasil.

[15] Jackson B. 2022. Website Downtime: Applicable Tips on How to Prevent It. Disponível em: <https://kinsta.com/blog/website-downtime/>. Acesso em: 24 set. 2022.

[16] Google. [s.d.]. Google Engineering Practices Documentation. Disponível em: <https://google.github.io/eng-practices/>. Acesso em: 28 jul. 2022.

[17] Zolkifli N.N.; Ngah A.; Deraman A. 2018. Version Control System: A Review. Procedia Computer Science 135: 408-415. DOI: 10.1016/j.procs.2018.08.191.

[18] Basques K. 2020. Por que HTTPS é importante. Disponível em: <https://web.dev/why-https-matters/>. Acesso em: 29 jul. 2022.

Como citar

Gonzaga F.L.; Bigaton A. Gerenciamento de risco para site de instituto de educação. Revista E&S. 2023; 4: e20230069.


Sobre os autores

Fabio Luis Gonzaga, Instituto de Pesquisa e Educação Continuada em Economia e Gestão de Empresas – Pecege – Tecnologia da Informação – R. Cezira Giovanoni Moretti, 580 – Santa Rosa – CEP 13414-157 – Piracicaba/SP, Brasil.

Aline Bigaton, Professora orientadora, Pecege – R. Cezira Giovanoni Moretti, 580 – Santa Rosa – CEP 13414-157 – Piracicaba/SP, Brasil.

Download link: PDF

You may also like