Working email and access to shared files alone do not tell you whether your IT infrastructure is healthy. Day-to-day use can conceal equipment whose maintenance coverage has expired, unnecessary administrator privileges, and backups that cannot be restored.
When reviewing your IT infrastructure, check its current operating condition, the potential impact of failures, and the requirements for recovery. The scope includes servers and network equipment, as well as cloud services, accounts, configuration information, maintenance contracts, and operating procedures.
1. Configuration Management: Understand Dependencies Between Business Operations and Systems
An IT asset inventory should record not only equipment names and quantities, but also the business operations each asset supports. Setting priorities during an incident requires understanding how failed equipment affects business operations.
For example, even if an order management system is running, users may be unable to log in if its authentication server fails. Cloud-based business systems also become inaccessible if the office internet connection or connectivity equipment fails.
Organize configuration information as follows.
| Managed Assets | Information to Record | Purpose of the Review |
| Servers and virtual environments | Purpose, OS, hosted systems, and connected systems | Identify business operations affected by an outage |
| Networks | Connections, routers, switches, and network topology | Check communication paths and the scope of impact from failures |
| Cloud services | Departments using the service, administrators, authentication methods, and contract information | Clarify management responsibilities and conditions of use |
| Data | Storage locations, systems using the data, and backup scope | Identify gaps in the information that needs protection |
| Maintenance and operations | Responsible staff, service providers, scope of support, and contact details | Facilitate communication and service requests during incidents |
Use configuration diagrams to identify locations where a failure of a single device or connection would make the entire system unavailable. These are known as single points of failure.
Even with two devices, both may fail at the same time if they share a power supply or connect to the same switch. Assess redundancy by the failures the configuration can withstand, not simply by the number of devices.
2. Lifecycle Management: Check Equipment Age and Support Status
Manage equipment age, hardware maintenance expiration dates, and OS and software support deadlines separately.
A server may still be covered by hardware maintenance even after support for its OS has ended. Conversely, even with up-to-date software, recovery can take longer if replacement parts for the equipment are unavailable.
Check the following during your review.
- Equipment installation dates and maintenance end dates
- OS, software, and firmware versions and support deadlines
- Patch installation status
- Availability of replacement parts or substitute equipment
- Support hours, scope of service, and excluded work under maintenance contracts
The support hours or on-site response times stated in a maintenance contract do not necessarily guarantee when business operations will be restored. Configuration and data may still need to be restored after parts are replaced.
Set replacement priorities based on equipment age, the business impact of an outage, exposure to external access, and available alternatives. If immediate replacement is not possible, specify interim measures such as limiting connectivity, together with a replacement deadline.

3. Access Management: Align Accounts and Permissions with Business Needs
Accounts belonging to former employees and permissions from previous roles may not cause visible problems in routine work. However, they leave unnecessary access paths open and increase the potential impact of account misuse.
Start by comparing current employees and their responsibilities with the accounts and permissions in each system.
In particular, treat administrator privileges separately from the permissions needed for everyday work. Using administrator accounts for routine email and web browsing can expose a broad range of systems if those credentials are compromised.
Include not only employee accounts but also maintenance accounts used by service providers and accounts used for system integration. For accounts with unknown owners, purposes, or expiration dates, check dependent processes before taking action.
Also check whether multifactor authentication is applied to administrator accounts and accounts used for external access, and whether activity logs are retained. Store emergency credentials in a way that controls access and records usage, without relying on an individual staff member's memory or email.
4. Backups: Verify Backup Results and Restore Capabilities
A successful backup job does not necessarily contain all the data you need. Backup settings may not have been updated when a new shared folder or business system was added.
First, compare the locations of data used in business operations with the backup inventory. Then check retention periods, backup frequency, storage locations, and restore methods.
Use the following two metrics when defining recovery requirements.
- RTO (Recovery Time Objective): The target time for restoring a system after an outage.
- RPO (Recovery Point Objective): The target point in time before an incident to which data must be restored.
Distinguishing these two metrics helps define measures for resuming operations quickly and measures for limiting data loss.
For example, if a business process cannot tolerate losing more than one hour of data, one backup per day cannot meet that requirement. Confirm the point to which data can actually be restored, considering backup intervals and any failed or delayed jobs.
In a restore test, verify not only that files can be retrieved but also that the required systems start and users can resume work. Missing configuration information, authentication, licenses, or connections can prevent business operations from resuming even after data is restored. Redundancy and synchronization serve different purposes from backups. If accidental deletions or corruption are replicated to a synchronized destination, a separate mechanism is needed to return to a previous healthy state. Also check that backup storage and administrative privileges are separated so that the backups are not affected by the same incident. Recovery time includes not only repairs and data restoration, but also detecting problems, assessing the situation, notifying stakeholders, and deciding what action to take. If monitoring alerts reach only one person, the maintenance contract number is unknown, or no one has been designated to authorize a shutdown or failover, the response may stall before technical work begins. Include the following information in incident response procedures. Pay attention to where procedures are stored. If instructions for shutting down an internal server are stored only on that server, they may be inaccessible during an incident. Provide a location that is less vulnerable to the incident and can be accessed securely by the staff who need it. Once the procedures are in place, check whether someone other than the primary contact can follow them. A contact list alone does not establish whether the necessary permissions and information are available. For each identified issue, record the affected system, business impact, interim measures, responsible person, and resolution deadline. Instead of leaving an item marked for review, specify what must be investigated to make a decision so that the next action is clear. Prioritize areas that directly support critical operations, have no alternative, and lack a verified recovery method. Some issues can be resolved by reviewing permissions or sharing contact information; others require equipment replacement or configuration changes. Start by selecting one business operation whose interruption would have a significant impact. Trace the systems, data, authentication, and networks it uses, then check its backups and recovery procedures. This turns gaps that an equipment inventory alone would not reveal into concrete improvements.
5. Recovery Readiness: Enable Response Even When the Primary Contact Is Unavailable
Turn Review Findings into an Improvement Plan



