XCP-ng
    • Categories
    • Recent
    • Tags
    • Popular
    • Users
    • Groups
    • Register
    • Login
    1. Home
    2. Popular
    Log in to post
    • All Time
    • Day
    • Week
    • Month
    • All Topics
    • New Topics
    • Watched Topics
    • Unreplied Topics

    • All categories
    • A

      VDI export to VMDK results in a corrupted disk

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Xen Orchestra
      11
      0 Votes
      11 Posts
      78 Views
      A
      I´m not a developer so I asked Ai what it "thinks" about it... ┌─────────────┐ │ XCP-ng VDI │ └──────┬──────┘ │ │ VDI_exportContent() │ format = VDI_FORMAT_VHD ▼ ┌──────────────────┐ │ VHD stream │ └────────┬─────────┘ │ │ vhdToVMDK() ▼ ┌─────────────────────────┐ │ vhdToVMDKIterator() │ └───────────┬─────────────┘ │ │ parseVhdToBlocks() ▼ ┌─────────────────────────┐ │ parseVhdStream() │ │ │ │ VHD blocks │ │ { id, data } │ └───────────┬─────────────┘ │ │ onlyBlocks() │ │ lba = id * blockSize ▼ ══════════════════════════════════════════ POSSIBLE ERROR HERE Incorrect mapping of id ↔ data or incorrect block ordering ══════════════════════════════════════════ │ ▼ ┌─────────────────────────┐ │ generateVmdkData() │ │ │ │ VHD block │ │ ↓ │ │ grainData │ │ ↓ │ │ createMarkedGrain() │ │ ↓ │ │ tableBuffer │ └───────────┬─────────────┘ │ ▼ ┌──────────────────┐ │ VMDK stream │ └────────┬─────────┘ │ ▼ ┌──────────────────┐ │ VMDK file │ └──────────────────┘
    • A

      Error: Can't init vhd directory without using alias

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Backup
      5
      1 Votes
      5 Posts
      84 Views
      A
      In any case, rolling back XO to an older build helps, so hopefully it will be fixed soon.
    • acebmxerA

      XCP Pulse collects XCP-ng and Xen Orchestra logs

      Watching Ignoring Scheduled Pinned Locked Moved Xen Orchestra
      4
      3
      3 Votes
      4 Posts
      147 Views
      acebmxerA
      v0.8.0–v0.9.1 — multi-user accounts, self-update, built-in HTTPS, and date ranges. Three releases since the last update, so bundling them here. Date ranges. Collections, extractions, findings runs and the support package can now be scoped to a window instead of always covering the whole bundle — last 24h/7d/30d, since last reboot, or a custom range. XO's own routes still have no date filter, so the first download is unchanged size — the range only narrows what gets kept and reported afterwards. Redact on demand. A collection no longer has to redact immediately with whatever rules happen to be on at that moment. There's now a checkbox to store the raw bundle only, and a "Redact now" button on the card afterwards using whichever rules are switched on when you press it. Multiple user accounts, roles, and an activity log. This was on the "what is coming" list in the first post and it's in now. Any number of accounts, three roles — admin, operator, viewer (read-only, but can still download what's already stored) — and an Activity page logging logins, settings changes, and job runs. Optional 2FA. Per-user TOTP, off by default, each account turns it on for itself. QR code setup, backup codes. Self-update from the UI, opt-in and off by default. Checking needs nothing extra; applying an update needs the Docker socket mounted in explicitly, since that's effectively host root, so it's a separate deliberate step — uncomment the socket mount and group_add block in docker-compose.yml, and it needs its own .env file (just DOCKER_GID=..., next to docker-compose.yml, not the same file as xcp-pulse.env) or docker compose up -d fails with unable to find group. .env.example has the one-liner to generate it. Built-in HTTPS, no reverse proxy required. Bundles nginx into the image to terminate TLS — self-signed cert on first start, or upload your own from Settings. A reverse proxy still works fine too. A user manual in the app itself, under /help — no need to leave XCP Pulse or check out the repo to read how something works. Fixed The account-requirements note from the first two posts was wrong in a way I only found by testing it properly: export:logs — the privilege log downloads actually need — isn't in any built-in XO role, but it can be granted through a custom role. "Test connection" was reading a catalogue that can't tell you that, and told every restricted account it needed full admin regardless. It now actually probes the download endpoint, and the docs walk through creating the custom role via three REST API calls (no UI for it yet on either XO version). A real concurrency bug: the shared SQLite connection wasn't locked across fetches, only execute, so two requests landing close together could crash a page reading jobs — reproduced it under the test suite's own concurrency test. Base-OS CVEs patched in the image build (perl-base, libc6, others); CI now runs a Trivy scan on every push. Still not fixed: the truncated-download-behind-a-reverse-proxy issue from the last post. Haven't found the actual cause yet. Also looking for anyone who can test against a remote proxy in a lab setup. Also open to any other suggestions, features, improvements, UI changes, etc... If any chance someone on Vates would test on their own time. Things like the support bundle and or the logs themselves are not being manipulated in any unwanted ways (for Vates or the project) https://github.com/acebmxer/xcp_pulse
    • C

      MS-01 performance issues w/ Intel 226 NICs

      Watching Ignoring Scheduled Pinned Locked Moved Hardware
      11
      0 Votes
      11 Posts
      3k Views
      B
      @Andrew said: pcie_aspm=disable Just to help anyone who would run into it. Disabling ASPM via dom0 settings/kernel did not resolve the issue. Had to disable it in BIOS (NUC13) After kernel level disable it did show: lspci -vv -s 55:00.0 | grep -E 'LnkCap|LnkCtl|LnkSta' LnkCap: Port #0, Speed 5GT/s, Width x1, ASPM L1, Exit Latency L0s <2us, L1 <4us LnkCtl: ASPM L1 Enabled; RCB 64 bytes Disabled- CommClk+ LnkSta: Speed 5GT/s, Width x1, TrErr- Train- SlotClk+ DLActive- BWMgmt- ABWMgmt- Only after BIOS disable it showed lspci -vv -s 55:00.0 | grep -E 'LnkCap|LnkCtl|LnkSta' LnkCap: Port #0, Speed 5GT/s, Width x1, ASPM L1, Exit Latency L0s <2us, L1 <4us LnkCtl: ASPM Disabled; RCB 64 bytes Disabled- CommClk+ LnkSta: Speed 5GT/s, Width x1, TrErr- Train- SlotClk+ DLActive- BWMgmt- ABWMgmt- LnkCtl2: Target Link Speed: 5GT/s, EnterCompliance- SpeedDis- LnkSta2: Current De-emphasis Level: -6dB, EqualizationComplete-, EqualizationPhase1- Running on NUC13ANBi7 Hope it helps anyone running into this problem.
    • BytevenidosB

      XO NFS option sec=krb5p encrypted transport

      Watching Ignoring Scheduled Pinned Locked Moved Xen Orchestra
      2
      0 Votes
      2 Posts
      42 Views
      poddingueP
      I gave this a go on my lab XOA, and XO isn't touching your option at all: the failure quotes the command it ran, mount -o sec=krb5p -t nfs <server>:/path /run/xo-server/mounts/<remote-id>, and a control remote with no custom options mounted the same export fine. It does fail with mount.nfs: an incorrect mount option was specified, but that line tells you less than it looks like: I fed it sec=totalnonsense and got the identical message back, which is why it reads like XO refusing a valid option when it's really just relaying what mount.nfs said. The appliance ships rpc.gssd as part of nfs-common, so that part's there. What isn't there is /etc/krb5.keytab or /etc/krb5.conf, and the systemd unit carries ConditionPathExists=/etc/krb5.keytab, so the daemon never starts. Mine last failed that condition at boot eleven days ago and said nothing about it. I went one step further: dropping a keytab in place is enough for rpc.gssd to start and stay up, so that condition really is the only thing stopping it, and it doesn't need krb5.conf for that. Whether the mount then works needs a KDC and principals that agree with each other, and I couldn't get that far, so that part is still untested.
    • K

      Intermittent Xen blkfront I/O stalls: all guest tags busy while tapdisk reports zero outstanding requests

      Watching Ignoring Scheduled Pinned Locked Moved Unsolved Compute
      17
      0 Votes
      17 Posts
      1k Views
      M
      Hello @anthoineb, Following your suggestion to test a newer blktap release, we installed blktap-3.55.5-9.4.xcpng8.3.x86_64 and matching debuginfo on all three hypervisors. We then fully shut down and started all 21 OpenSearch data VMs, completing this on September 11. We verified that their running tapdisk executables matched the installed binary. Unfortunately, the stall recurred on September 13 on os-ott-data-1-5. For the affected tapdisk process (PID 2797568), we verified during the incident: /proc/2797568/exe pointed to /usr/libexec/tapdisk, without (deleted); its SHA-256 matched the installed executable; GDB loaded matching debug symbols, build ID dc98a78aff623c1cf12518663c859e3efa32413c. Before recovery: guest I/O made no progress; 255 requests were in flight and 256/256 tags were busy; I/O PSI full was approximately 96–97%, with four tasks in D state; tapdisk reported zero outstanding requests. The active td_xenblkif ring showed: req_prod = 866947066 req_cons = 866946810 rsp_prod = 866946810 rsp_prod_pvt = 866946810 nr_ents = 256 Thus, 256 requests were pending in the active PV ring but had not been consumed. After preserving the baseline GDB capture, our recovery controller executed your suggested call once: call (void)tapdisk_xenblkif_sched_chkrng(blkif) The call completed at 09:40:14 MSK (UTC+03:00). By 09:40:25, guest I/O was progressing again, inflight requests and busy tags were zero, and the local OpenSearch API was responding. The post-recovery GDB capture showed all four ring counters equal to 866967322. The node rejoined the cluster at 09:42:00 without restarting OpenSearch or rebooting the VM. So the same ring-processing stall still occurs with the running 9.4 binary, and the explicit ring-check call still restores I/O. We have complete before/after GDB captures and the recovery-call log available. Is there a newer build or a specific scheduler/event-channel state you would like us to capture during the next occurrence?