Documentation PI Nexus+ Documentation

PI Nexus+ / Reference

Host Monitoring Reference

This reference lists the issues, limits, rights, ports and timings of host monitoring. The Administration Guide and the User Guide link here instead of repeating these values.

Overview

This reference lists the issues, limits, rights, ports and timings of host monitoring. The Administration Guide and the User Guide link here instead of repeating these values.

Host issues

Host issues appear in the Issues page list, marked Hosts (the Category filter has no Hosts entry), on the Issues tab of the host dialog, and on the sensor named in the last column. Values marked adjustable are set under Admin > Hosts > Limits or per host (see Default limits).

IssueSeverityRaised whenResolves whenSensorTypical fix
Low disk spaceWarning below 10 % free, Error below 5 % (adjustable; per disk also in GB)A disk is below its limit on a pollThe disk is above the limit againDisk and the drive, for example Disk C:Free space or extend the disk; set a disk's own limit for large archive disks
Disk filling upWarningAt the trend of the last 7 days, the disk is full within 30 days (adjustable, 0 = off). Needs at least 24 hours of data. Not raised while Low disk space is openThe forecast is beyond the limitDisk and the driveFind what is growing; plan more space
High CPUWarning; optional Error. Off by defaultThe average CPU over the window (default 30 minutes) is above the limit; one missed poll is toleratedThe average is more than 5 points below the Warning limitCPUFind the busy process on the Processes section
High memory useWarning; optional Error. Per host only, off by defaultThe average memory use over the window is above the limitAs for High CPUMemoryCheck the processes; do not switch this on for PI Data Archive or SQL Server hosts, which use memory as cache
Host not answeringWarningThree polls in a row failed. Disk, CPU and memory issues keep their state meanwhileA poll succeedsHostCheck that the host runs and that WinRM answers; run Test access
Service not answeringErrorAn availability check failed 3 times in a row (adjustable 1 to 10)2 good checks in a row (adjustable 1 to 10)The check, for example PI Data Archive PISRV01Check the PI service on the host and the network path from the scanner server
Certificate expiringWarning within 30 days; Error when expired or invalidAn HTTPS check's certificateThe certificate is renewedThe HTTPS checkRenew the certificate
Host restartedWarningThe host's boot time moved by more than 2 minutes. Not raised inside a maintenance windowAutomatically after 24 hoursUptimeNone needed if planned; otherwise check the Windows event log
Clock driftWarning above 5 s, Error above 60 s (adjustable)The host's clock differs from the scanner server's by more than the limit plus the measurement's uncertaintyThe drift is within the limitClockFix the host's time synchronization
Process not runningErrorA watched process is missing in as many polls in a row as Not answering after failed checks (default 3)The process runs againProcess · and the program nameStart the PI component; tick Forget the role processes and application pools seen so far after removing one on purpose
Windows service stoppedErrorA watched service with start mode Automatic is not running in 2 polls in a row. The message names the state and exit codeThe service runsThe service name (group Services)Start the service, or set an unused service to Manual or Disabled on the host
Windows service restartingWarningA watched service's process changed 3 times within 24 hours without a host restartNo further restarts in 24 hoursThe service nameCheck the service's own log and the Windows event log
Windows service disabledWarningA watched service was switched from Automatic to Disabled while monitoredAutomatically after 24 hoursThe service nameNone needed if intended
Slow diskWarning. Off by defaultA disk's average access time (the slower of read and write) is above the limit over 30 minutesThe average is more than 5 ms below the limitDisk latencyCheck the storage and the load on the disk
Network errorsWarning. Off by defaultA network interface reported errors in every poll of 30 minutesA poll of the last 30 minutes had no errorsNetworkCheck the NIC, driver, cable and switch port
PI Backup failedErrorThe PI Backup Subsystem reports that its last backup failedThe next backup succeedsPI BackupCheck the PI Data Archive message log and the backup target
PI event queue growingWarningThe PI Snapshot event queue grew in every poll of 30 minutesThe queue stops growingEvent queueCheck the PI Archive Subsystem and disk performance
PI Buffer queue growingWarning; Error above a size (adjustable, off by default)The buffer queue of an interface node for one PI Data Archive grew in every poll of 30 minutesThe queue stops growingBuffer · and the PI Data ArchiveCheck the connection from the node to the PI Data Archive
SQL backup oldWarningA database has had no full backup for 2 days (adjustable, 0 = off). See the note belowA full backup is takenBackup · PI System or Backup · otherFix the SQL Server backup job
SQL log fillingWarning above 80 %, Error above 95 % used (adjustable)A database's transaction log is above the limitThe log is below the limitLog · and the databaseBack up the log (FULL recovery) or check long-running transactions
SQL blockingWarning. Off by defaultA session has been blocked longer than the limit (5 minutes when switched on)No session is blocked longer than the limitBlockingFind the blocking session in SQL Server
Application pool stoppedErrorAn IIS application pool that was seen running is stoppedThe pool runsApp pool · and the pool nameStart the pool in IIS Manager

Note: SQL backup old covers online databases in FULL or SIMPLE recovery, except tempdb and snapshots. A database never backed up counts from its creation. Each SQL Server instance has at most two such issues: one for the PI System's databases (the AF database PIFD, and the PI Vision and PI Nexus+ databases PI Nexus+ connects to) and one for all others. Each names the three oldest.

Pausing or deleting a host resolves its issues without an email. A disk marked PI archive volume is named "PI archive volume" in its issues.

Monitoring board statuses

A host has the status of its worst sensor. The counts at the top of the board count sensors.

StatusMeaning
DownThe host failed 3 polls in a row, an availability check reached its failed-checks limit, or a PI Interface or PI Adapter on the host reports Error
ErrorAn open Error issue on the sensor
WarningAn open Warning issue; a missed check or one or two missed polls; or an interface or adapter that reports Warning or Unknown
StaleNo new data for 3 intervals: 15 minutes for a polled host; for a host with availability checks only, 3 check intervals and at least 5 minutes
UpEverything answers
MaintenanceThe host is inside a maintenance window. Never counted as a problem
PausedMonitoring is paused or access is not confirmed. Interfaces and adapters on the host are still monitored and can change the host's status

Suppressed issues do not change a status, except that a host or check that does not answer is Down even when its issue is suppressed. Several status counts can be combined as a filter. Hosts in maintenance and paused hosts have no rows on the Alarms tab. On the Capacity tab, a disk that is not filling shows Not filling.

Host status on Servers > Hosts

StatusMeaning
MonitoredThe last poll succeeded
Not answeringThe last poll failed
PausedMonitoring is paused
No accessThe access test has not passed
Waiting for first pollAccess passed; no poll yet

Sensors per group

GroupSensors
ResourcesCPU, Memory, Disk per drive, Clock, Uptime, Network, Disk latency, and Process · name while a watched process is missing
AvailabilityHost (polled hosts) and one sensor per availability check, for example PI Data Archive PISRV01 or TCP port 1433
PI Data ArchiveEvent queue, PI Backup, Archive fill
PI BufferBuffer · and each PI Data Archive the node buffers to
SQL ServerBackup · PI System, Backup · other, Log · and each database, Blocking
PI VisionApp pool · and each IIS application pool
ServicesEach watched Windows service
InterfacesEach PI Interface with runtime monitoring on the host
AdaptersEach PI Adapter service on the host (the worst of its Collectors)

A sensor the host does not deliver is left out. A host read from PI points has only CPU, memory and disk sensors, plus its availability checks.

Role readings

RoleRead in each poll
PI Data ArchiveArchived, snapshot and out-of-order events per second; PI Snapshot event queue; primary archive fill; PI Backup result (failed or not; the counters carry no backup time)
Interface nodePer PI Data Archive buffered to: queued events, events sent per second, connection; round trip from the node to up to 3 PI Data Archives (a ping run on the node)
SQL Server (counters)Page life expectancy, batch requests per second, blocked processes, memory grants pending, per instance including named instances
SQL Server (T-SQL)Per database: data and log size, growth, log used, age of the last full and log backup; blocked sessions. Only on PI Nexus+'s own SQL Server, the PI Vision SQL Server, and instances added under SQL Server instances
PI VisionState and recycles of each IIS application pool; connections and requests per second per web site; w3wp memory

Default limits

Set the defaults under Admin > Hosts > Limits. A host overrides Disk, CPU and Disk latency by clearing Use defaults in its Edit dialog; a disk overrides the host under Limits > Volumes. Disk settings apply in the order disk, host, default; Ignore this volume wins over all.

SettingDefaultRangeSet in
Disk Warning below % free101 to 99Limits, host, disk
Disk Error below % free51 to 99, below WarningLimits, host, disk
Disk Warning/Error below GB freeNot setAbove 0 to 100,000 GBDisk only
Disk Forecast days (0 = off)300 to 365Limits, host, disk
Alert on sustained CPUOffLimits, host
CPU Warning above %90 when switched on1 to 99Limits, host
CPU Error above %Off1 to 99, above WarningLimits, host
CPU Sustained minutes3010 to 240, steps of 5Limits, host
Alert on sustained memory useOffSuggested 95 % over 30 minutesHost only
Check every1 minute1 or 5 minutesLimits
Not answering after failed checks31 to 10Limits
Recovered after good checks21 to 10Limits
Clock drift Warning above seconds51 to 3600Limits
Clock drift Error above seconds601 to 3600, above WarningLimits
Alert on sustained disk latencyOff; 50 ms suggested1 to 10,000 msLimits, host
Alert on network errorsOffLimits
Error when a PI Buffer queue is above a sizeOff; 100,000 events suggested1 to 1,000,000,000Limits
Full backup older than days (0 = off)20 to 365Limits
Log Warning above % used801 to 99Limits
Log Error above % used951 to 100, above WarningLimits
Alert on blocked sessionsOff; 5 minutes when on1 to 1440 minutesLimits

These values are fixed: Host not answering after 3 failed polls; Certificate expiring 30 days ahead; Windows service stopped after 2 polls; Windows service restarting at 3 restarts in 24 hours; 30-minute windows for Slow disk, Network errors, PI event queue growing and PI Buffer queue growing. A changed limit applies from the next poll or check; past data is not re-evaluated.

Intervals and timings

WhatInterval or time
Resource poll (CPU, memory, disks, services, roles)Every 5 minutes
Availability checksEvery minute (or every 5 minutes)
Access test after adding or editing a hostWithin seconds; 3 hosts at a time
Disk forecast recalculationEvery hour
Discovery of role counters (PI, SQL Server, IIS)At the access test and once a day
Monitoring page refreshEvery minute while visible
Red "data is old" banner on the Monitoring pageData older than 15 minutes
No Windows service findings after a host startsFirst 15 minutes
PI point value counted as missedOlder than 15 minutes or bad quality
Host restarted and Windows service disabled stay open24 hours
Issues opened during a maintenance window are emailedWithin a minute after the window ends
Scanner service watchdog emailAfter 10 minutes without a report from the Scanner service (checked every minute)
Reminder and escalation emailsChecked with every availability check cycle

Timeouts

OperationTimeout
WinRM port test5 s
WinRM (CIM) operation20 s
One host's poll60 s
HTTP check, PI Data Archive or AF Server check15 s
TCP check5 s
SQL Server connect and command5 s; 10 s per instance in total

Retention

DataKept for
Five-minute samples and individual check results14 days
Hourly averages (with lowest and highest value)400 days
Restart markers400 days

Ports

All connections start from the scanner server unless stated otherwise.

ToPortUsed for
Monitored hostTCP 5985WinRM over HTTP (default)
Monitored hostTCP 5986WinRM over HTTPS (credential profile)
PI Data ArchiveTCP 5450Availability check; PI points hosts
AF ServerTCP 5457Availability check
PI VisionTCP 443 or 80Availability check of the PI Vision URL
SQL ServerTCP 1433 or the instance portSQL Server reads over T-SQL (optional)
AnyAs configuredAvailability checks you add
From an interface node to its PI Data ArchivesICMP echoPI Buffer round trip (shown only; Not available where blocked)

Required Windows rights on each host

Grant these to the scanner service account, or to the account of the host's credential profile. No administrator rights are needed.

RightWithout it
Member of Remote Management UsersWinRM refuses the sign-in
Enable Account and Remote Enable on the WMI namespace Root\CIMV2Every reading is denied
Member of Performance Monitor UsersCPU, disk latency, network, process and role counters return no data; memory and disks still work
Read access to the Service Control Manager and to each PI-related service (optional; granted by the rights script)Windows services are not read; PI processes are watched instead

WinRM must be enabled on the host (winrm quickconfig).

Rights script parameters

Grant-PINexusHostAccess.ps1 runs in Windows PowerShell 5.1 (not PowerShell 7) as a local administrator.

ParameterMeaning
-AccountRequired. The account to grant, as DOMAIN\user, HOST\user or user@domain
-ServiceFurther service names to grant read access to, comma-separated, for example 'MyOpcServer,OtherService'
-SkipServicesLeaves out the Windows services step
-WhatIfShows what would change without changing it

The script finds the local groups by their well-known SIDs, so it works on any Windows language. It changes only what is missing and can be run again. It runs Enable-PSRemoting -SkipNetworkProfileCheck only when WinRM has no running listener. It prints the security descriptor of each service it changes, before and after; to undo one grant, run sc.exe sdset <service> "<before>".

The script refuses to run on a domain controller, which has no local groups: it stops before any step, also with -WhatIf, and changes nothing. Grant the rights on a domain controller through domain policy instead.

Group Policy

For many hosts, grant the same rights with a GPO linked to the servers' OU.

WhatPolicy setting
Local groupsComputer Configuration > Preferences > Control Panel Settings > Local Users and Groups: one Group item for Remote Management Users (built-in) and one for Performance Monitor Users (built-in), action Update, Add the scanner service account. Restricted Groups also works, but replaces the whole membership
WinRM listenerComputer Configuration > Policies > Administrative Templates > Windows Components > Windows Remote Management (WinRM) > WinRM Service > Allow remote server management through WinRM: Enabled, IPv4 filter * or the scanner server's address
WinRM serviceComputer Configuration > Policies > Windows Settings > Security Settings > System Services > Windows Remote Management (WS-Management): Automatic
FirewallWindows Defender Firewall with Advanced Security, inbound rule Windows Remote Management (HTTP-In), TCP 5985, ideally limited to the scanner server
WMI namespace and service rightsNo policy setting. Run Grant-PINexusHostAccess.ps1 -Account 'DOMAIN\svc-pinexus' as a computer startup script (Computer Configuration > Policies > Windows Settings > Scripts (Startup/Shutdown) > Startup, PowerShell Scripts tab, the account as script parameter), or as an immediate scheduled task running as SYSTEM (Preferences > Control Panel Settings > Scheduled Tasks). The script changes only what is missing, so it can run at every start

After gpupdate /force (or the next policy refresh), restart hosts where WinRM was newly enabled, then choose Test access. Group membership applies to new sign-ins only.

Required SQL Server rights (optional)

The scanner service account reads SQL Server with Windows authentication. All reads are SELECT statements and DBCC SQLPERF(LOGSPACE); nothing is written.

RightGrant withWithout it
A loginCREATE LOGIN [DOMAIN\svc-pinexus] FROM WINDOWS;Nothing is read over T-SQL; counters still are
VIEW SERVER STATEGRANT VIEW SERVER STATE TO [DOMAIN\svc-pinexus];Log use and blocked sessions are not read
VIEW ANY DEFINITIONGRANT VIEW ANY DEFINITION TO [DOMAIN\svc-pinexus];File sizes and growth are not read
SELECT on msdb.dbo.backupsetIn msdb: CREATE USER [DOMAIN\svc-pinexus] FOR LOGIN [DOMAIN\svc-pinexus]; GRANT SELECT ON dbo.backupset TO [DOMAIN\svc-pinexus];Backup age is not read

Access test steps

StepRequiredChecks
Name lookupYesThe host name resolves from the scanner server
WinRM portYesTCP 5985 or 5986 answers
WinRM sign-inYesThe account can sign in; names the account and transport
WMI accessYesRoot\CIMV2 can be read
Memory, CPU counters, DisksYesThe core readings
PI pointsYes (PI points hosts only)The mapped PI points can be read
Disk latency (optional), Network (optional), Processes (optional)NoPerformance counters
Windows services (optional)NoHow many services the account can read and how many are watched
PI Data Archive counters (optional), PI Buffer counters (optional), SQL Server counters (optional), SQL Server databases over T-SQL (optional), IIS counters (optional)NoRole readings; listed only for hosts with that role

An optional step that is not available never raises an issue; the host is monitored without it.

Access problems

Open the host's Access details (the Access label, or Access details in the row menu) and start with the first failed step.

Step or symptomCauseFix
Name lookup failedThe name does not resolve on the scanner serverCorrect the address or the DNS entry
WinRM port failedWinRM is off, or a firewall blocks TCP 5985 or 5986Run winrm quickconfig (HTTPS: create the HTTPS listener) and open the port from the scanner server
WinRM sign-in failed: logon rejectedWrong account or password, or the account is disabled, locked or expiredCheck the account; for a profile, enter the password again under Credentials
WinRM sign-in failed: access deniedThe account is not in Remote Management UsersRun the rights script
WinRM sign-in failed: IP addressAn IP address with the scanner service account or over HTTPAdd the host by DNS name, or use a credential profile over HTTPS
WinRM sign-in failed: certificate refusedThe HTTPS certificate is not trusted or not issued to the host's nameIssue a certificate for the host's name from a trusted CA, or use Skip CA check / Skip name check
WinRM sign-in failed: password cannot be decryptedSplit installation without the shared secret certificate on both serversInstall the certificate on both servers, then enter the profile's password again
WMI access failedNo Enable Account / Remote Enable on Root\CIMV2Run the rights script
CPU counters failed, or an optional counter step not availableThe account is not in Performance Monitor UsersRun the rights script; sign-ins made before the change do not see it, so test again
Windows services (optional) not availableService rights not granted, or a PI component was installed after the script ranRun the rights script again on the host
SQL Server databases over T-SQL (optional) not availableNo login, or missing rights on the instanceGrant the rights in Required SQL Server rights (optional)
PI points failedThe PI Data Archive is not reachable, or the points are wrongCheck the PIPerfMon points and use Find points
Host shows Stale on the boardThe Scanner service has stoppedStart the PI Nexus+ Scanner service
Disk issue on a disk that is fine with little free spacePercent limits do not suit a large diskSet GB free limits for that disk
Windows service stopped for a service you do not useThe service is set to AutomaticSet it to Manual or Disabled on the host, or add it to Never watch

The host's row menu also has Test access, Edit, Pause monitoring and Delete.

Watched Windows services

Watched services raise the three Windows service issues. Only a service with start mode Automatic is expected to run.

Role or ruleServices
PI Data Archivepiarchss, pibackup, pibasess, pilicmgr, pimsgss, pinetmgr, pishutev, pisnapss, pisqlss, pitotal, piupdmgr
AF ServerAFService, PIAnalysisManager, PINotificationsService
SQL ServerMSSQLSERVER, MSSQL$instance, SQLSERVERAGENT, SQLAgent$instance
PI VisionW3SVC, WAS
Interface nodepibufss
Any hostEvery service whose program is installed under \PI\bin\, \PI\adm\, \PIPC\, \OSIsoft\ or \AVEVA\ (each PI interface and PI Adapter instance)
Added by youUp to 25 names under Windows services > Also watch; Never watch excludes a service and wins

pishutev stopping cleanly after startup is expected. A service that is uninstalled is not reported as stopped. A process whose service is watched is reported as that service, not as Process not running.

Watched processes

Processes are watched where Windows services cannot be read. A role process is expected only after it has been seen running on that host.

RoleProcesses
PI Data Archivepiarchss, pisnapss, pibasess, pinetmgr, piupdmgr, pimsgss, pibackup
AF ServerAFService, PIAnalysisManager, PIAnalysisProcessor, PINotificationsService
SQL Serversqlservr
PI Visionw3wp
Interface nodepibufss
Adapter nodeOSIsoft.Data.System.Host
Added by youUp to 25 program names (without .exe) under Processes > Also watch; always expected

PI points read for a PI points host

Find points searches the chosen PI Data Archive for points whose instrument tag (or extended descriptor) is one of these counter paths with one of the host's names.

MetricCounter path
CPU\\HOST\Processor(_Total)\% Processor Time
Memory\\HOST\Memory\Available MBytes or \\HOST\Memory\% Committed Bytes In Use
Disk free %, per drive\\HOST\LogicalDisk(D:)\% Free Space
Disk free MB, per drive\\HOST\LogicalDisk(D:)\Free Megabytes

Memory shows as used percent only when the host's memory size is known from an earlier WinRM reading. A disk with only % Free Space has no size, no GB limit and no forecast. A host has at most 60 point mappings.

Availability checks

CheckCreatedHow it tests
PI Data Archive nameAutomatically for a PI Data Archive targetTCP 5450, then a server-time request. TCP only for OpenID Connect targets
AF Server nameAutomatically for an AF Server targetTCP 5457, then a server-time request. TCP only for OpenID Connect targets
PI Vision nameAutomatically for a PI Vision instanceOpens the URL with Windows authentication; at most 3 redirects on the same host
HTTP addressBy youOpens the address. Sends Windows credentials only when the address names the host, or with Send Windows credentials
TCP portBy youConnects to the port (1 to 65535)

Checks run from the scanner server, also while the host's WinRM access fails. In the Availability checks dialog, Switch off stops a check and Delete removes one you added.

Notifications

SettingWhereDefault
Host monitoring delivery ruleAdmin > Notifications > Delivery RulesOff
New issues, Severity escalations, Recoveries after an alertSame ruleOn once the rule is on
Minimum severitySame ruleWarnings and errors (or Errors only)
ReminderSame ruleOff; every 1, 2, 4, 8, 12 or 24 hours
EscalationSame ruleOff; after 15 or 30 minutes, or 1, 2 or 4 hours
Host monitoring email typeEach recipientTicked
Host monitoring escalations email typeEach recipientNot ticked

Reminders and escalations apply only to open Error issues that are not acknowledged or suppressed and not inside a maintenance window. A reminder repeats after the chosen time since the last email about the issue; an escalation is sent once per occurrence. Maintenance windows suppress all host emails. The scanner service watchdog sends PI Nexus+ scanner service is not running and running again to the Host monitoring recipients when the rule is on.

Limits on size

ItemLimit
Hosts added in one Add N selected100
Hosts polled in parallel8
Availability checks run in parallel32
Processes or services added per host (Also watch, Never watch)25 each
SQL Server instances added per host10
Databases read per SQL Server instance200
PI point mappings per host60
Address of an HTTP check512 characters
Maintenance window5 minutes to 31 days; weekly windows up to 7 days
Maintenance note512 characters
Rows on the Alarms tab500 (Download CSV includes all)

Uptime figures

FigureMeaning
AvailabilityShare of good checks in the period, each hour weighted by how often it was checked. Hours that overlap a maintenance window are left out. Green at 100 %, amber from 99 %, red below
Average responseAverage time of the answered checks
OutagesRuns of failed checks as long as it takes to raise Service not answering (3 by default). A host's outages are its checks' outages merged
Longest outageLongest such run

Who can do what

ActionRole
View Monitoring and Servers > HostsViewer
Acknowledge or suppress host issues; Download CSVOperator
Everything under Admin > Hosts; host email settingsAdmin

Support bundle

The support bundle includes host-monitoring.json: hosts, access test results, last poll state, connection settings, checks, limits and credential profiles by name and account. Passwords, and user names or passwords in check addresses, are never included.