IT Service Manager ExaminationStudy guide
SM (IT Service Manager): Japan’s top-tier national certification for IT service operations. This course targets the multiple-choice morning exam, centered on the SM-specific Part-A-II specialty (service management fundamentals, design/transition, operation, continuity/availability, security operations, and system operation/facilities). The common Part-A-I builds on the AP course; the descriptive/essay afternoon exam is out of scope.
About IT Service Manager Examination (SM)
IT Service Manager Examination (SM) is a Professional / Expert-level certification from IPA(情報処理技術者試験). This page organizes the exam scope into a 6-chapter, 25-section study guide and lets you check your understanding with exam-style practice questions. A good flow is to read the chapters below in order, then test yourself via "Practice questions."
Exam domains (approximate weighting)
- Service management fundamentals & SMS~16%
- Service design & transition~20%
- Service operation~20%
- Service continuity & availability~16%
- Security operations & supplier management~14%
- System operation & facilities~14%
Weights are approximate guidance for the live exam. Each domain is covered in detail in the chapters and sections below.
Official exam information: https://www.ipa.go.jp/shiken/kubun/sm.html
1Service management fundamentals & SMS
- 1.1ITIL and the big picture of service management
Grasp how ITIL 4's Service Value System (SVS), the seven guiding principles, the four dimensions, and the Service Value Chain (SVC) mesh together to co-create value, from a service manager's decision-making viewpoint.
- 1.2JIS Q 20000 and the SMS
Understand the requirements of the service management system (SMS) defined by JIS Q 20000, and the roles of PDCA continual improvement and third-party certification, from the angle of designing operations that withstand audit.
- 1.3Service lifecycle, value co-creation, and SLM
Grasp that a service's value is not set by the provider alone but co-created with the consumer, and that across the service lifecycle, SLM (service level management) sustains and improves the agreed value—from an operational-judgment viewpoint.
- 1.4Service catalogue and the place of service level management
Understand how the service catalogue makes live services visible from the customer's view, and the place of service level management that agrees and sustains SLAs on top of it—framed by balancing customer expectations against cost.
2Service design & transition
- 2.1SLA/OLA/UC & service-level design
Covers the three-layer structure of the SLA agreed between provider and customer, the OLA between internal operations teams, and the UC (underpinning contract) with external suppliers, how SLM maintains and improves them, and how to design service-level targets amid trade-offs of availability, cost, and feasibility.
- 2.2Change management
Covers choosing among three change paths—standard changes that are pre-approved and routine, normal changes assessed by the CAB, and emergency changes decided by the ECAB and documented afterward—and how to make changes to a live service safely through risk assessment and a back-out (rollback) plan.
- 2.3Release & deployment management
Covers release and deployment management, which safely delivers approved changes to production: designing the release unit (what is packaged together), choosing a deployment approach such as phased rollout or a pilot, and how to plan distribution to a live service through the separation of build, test, and deploy.
- 2.4Configuration management
Covers configuration management, which accurately manages the CIs (configuration items) that make up a service in the CMDB, its linkage with change management and release management, and how it differs from IT asset management, whose main purpose is grasping owned assets—all from the practical angle of impact analysis and maintaining accuracy.
- 2.5Designing availability & capacity (at design time)
Covers how, at the design stage, to realize an availability target derived from requirements—redundant design that eliminates a single point of failure, series (product) and parallel (1−(1−a)ⁿ) availability, and designing capacity from a future demand forecast—all amid trade-offs with cost.
3Service operation
- 3.1Incident management
Centered on the fact that the purpose of incident management is the rapid restoration of service, not the investigation of root cause, this section covers prioritization by impact x urgency, escalation (functional and hierarchical), the separate handling of a major incident, and when to apply a workaround to restore service even before the root cause is known—together with how to judge these under SLA-target constraints.
- 3.2Problem management
Covers how the purpose of problem management is to prevent recurrence through root-cause identification and a permanent fix (in contrast with incident management, whose purpose is rapid restoration), the flow of raising a problem from recurring faults or major incidents and performing root-cause analysis (RCA), the use of a known error database (KEDB) that accumulates causes and workarounds, and choosing between reactive (waiting for a fault) and proactive (getting ahead via trend analysis) problem management.
- 3.3Event management and request fulfilment
Covers event management, which monitors state changes of components, catches signs that exceed a threshold, and can automatically raise incidents (distinguishing informational, warning, and exception events), and request fulfilment, which handles routine, low-risk service requests such as password resets or access grants through a flow separate from ordinary incidents—together with judging which report to route to which process.
- 3.4The service desk
Covers the role of the service desk as the single point of contact (SPOC) for users (centralizing communication and maintaining user satisfaction) and the structural patterns—local, centralized, virtual, and follow-the-sun—together with the judgment of which pattern to choose under constraints such as site distribution, language, the need for 24-hour coverage, and cost.
4Service continuity & availability
- 4.1IT service continuity management (ITSCM), BCP, and disaster recovery
Covers IT service continuity management (ITSCM), which restores a service within a target time after a major disruption such as a disaster; its relationship to the BCP (the business-wide continuity plan); the numeric recovery objectives RTO (recovery time objective), RPO (recovery point objective), and MTPD (maximum tolerable period of disruption); and the choice of recovery site among hot, warm, and cold sites, building judgment for selecting a continuity approach under constraints of business impact and cost.
- 4.2Availability management (availability rate, redundancy, single point of failure)
Covers the availability rate, computed as MTBF / (MTBF + MTTR); the distinction between the mean time between failures (reliability, MTBF) and the mean time to repair (maintainability, MTTR); estimating overall availability as the product of availabilities for a series configuration and as 1 - (1 - a)^n for parallel redundancy; and eliminating a single point of failure (SPOF) through redundancy.
- 4.3Capacity management (demand forecasting, thresholds, trend analysis)
Covers capacity management and its three sub-processes (business capacity management, service capacity management, and component capacity management), which secure exactly the performance and volume a service needs; demand management, which levels demand itself; early detection via threshold monitoring and trend analysis; and identifying the bottleneck, building judgment for choosing between augmentation and demand suppression.
- 4.4Backup and recovery (methods, RTO/RPO)
Covers the differences and recovery procedures among full backup, incremental backup, and differential backup; generation management that retains multiple generations; off-site storage as protection against disaster; and how backup methods correspond to the recovery objectives (RTO, RPO), building judgment for selecting a method from business requirements.
5Security operations & supplier management
- 5.1Security operations
Covers the operational controls that keep information security running in live service delivery-priority judgment for patch management and vulnerability management, log collection and monitoring, the SIEM that correlates multiple logs, and the handoff that connects detected incidents into incident management-framed as the judgments of an operator upholding SLA targets.
- 5.2Access management & privileged access
Covers access management, which controls who is allowed what in service operations-the principle of least privilege, the provisioning (granting) and periodic review (revocation) of privileges, privileged access management that handles powerful rights strictly, and the authentication that verifies identity-framed as an operator's judgment to prevent risks such as insider misuse and dangling leaver accounts.
- 5.3Supplier management & underpinning contracts
Covers supplier management, which aligns externally dependent parts with the SLA-the UC (underpinning contract) with external suppliers, aligning the customer SLA with UCs/OLAs, supplier evaluation and review, the demarcation of responsibility during faults, and the shared responsibility model when using the cloud-framed as the judgment of a service manager who ultimately upholds the SLA target.
- 5.4Service audit, corrective action & continual improvement
Covers the mechanisms that keep service management running and improving-the internal audit that detects nonconformities, the corrective action for findings (a permanent fix to the root cause, not a stopgap), CSI (continual service improvement), the PDCA cycle, and the KPIs/metrics that steer improvement by the numbers-framed as the judgment of a service manager improving the service based on measurements.
6System operation & facilities
- 6.1System monitoring and job management
Covers monitoring items (CPU, memory, disk, response time) and the design of thresholds and alerts that catch anomalies early, job scheduling that protects the SLA, the dependencies and abnormal-end handling (abort, rerun, skip) of a job net that ties multiple jobs together, and two-stage warning/critical thresholds that suppress false alarms.
- 6.2Backup operation, automation, and operating procedures
Covers operational automation that hands routine work to machines, the runbook / operating procedures that document steps, human-error prevention through double-checks and call-and-response, and safely applying configuration changes (change-management linkage and rollback) to production. The key is separating routine work to automate from exceptions requiring human judgment.
- 6.3Facility management
Covers the UPS (uninterruptible power supply) and on-site generator that bridge a momentary outage to standby power, the air conditioning that cools equipment and the power distribution that delivers electricity, seismic/base-isolation protection against earthquakes, access control preventing unauthorized entry, and the data-center Tier levels that stage a facility's availability. The key is choosing the degree of redundancy by trading off the availability target against cost.
- 6.4Operations design and cost, staffing, and budget management
Covers operations design that builds the operating structure and procedures, the design of shifts and staffing that runs a 24-hour operation seamlessly, cost optimization that balances investment and running cost, KPIs that measure operational quality numerically, and supplier / resource / budget management that governs external suppliers, resources, and budget. The key is building a structure that is neither excessive nor deficient within the SLA-versus-cost trade-off.

